anthropics / anthropics/claude-code
Response-level safety classifier false positives on authorized DeFi liquidation-keeper research (4 stops)
Nobody has claimed this yet.
- Dominant language
- TypeScript
- Stars
- 147k
- Forks
- 24k
- PR merge metrics
- PR metrics pending
Description
Environment: Claude Code desktop app (Code tab), Claude Sonnet 5, 2026-09-19/20.
Message received (4 times): "Your response above was stopped by a safety classifier — this is not a tool or API error. The rest of it was withheld, and tool calls in it that had not finished did not run. Do not produce that content again, even reworded." Tool calls showed "Interrupted: the response that made this tool call was stopped by a safety classifier while the call was running."
Work: Permissionless liquidation-keeper research using documented public liquidate entry points, flash-loan funded only. No exploit, no attack on any system.
What was stopped (all in one narrow area): two analytical/read-only scripts (reserve-config dump, borrower classification by oracle-feed freshness), a read-only pause-status check script, and a web search for a price-oracle contract address, all in a workflow about Pyth-based pull oracles that had gone stale. Dozens of other scripts and on-chain reads in the same session were never stopped.
Why I believe these are false positives: submitting a genuine, publisher-signed Pyth update is the documented permissionless refresh mechanism, not manipulation; liquidations only apply when the protocol's own oracle marks a position unhealthy.
Ask: please review for false positives, and point to any verification path for legitimate DeFi security researchers. I've also emailed usersafety@anthropic.com.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No repository files, tests, or implementation entry points are named. Start by locating the response-level safety-classifier handling in Claude Code and determine whether this report maps to an actionable product or policy change; done requires a maintainer-defined verification path or confirmed classifier fix.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, security
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100