anthropics / anthropics/claude-code

Response-level safety classifier false positives on authorized DeFi liquidation-keeper research (4 stops)

Open
#95,647 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

area:desktop area:model bug
Dominant language
TypeScript
Stars
147k
Forks
24k
PR merge metrics
PR metrics pending

Description

Environment: Claude Code desktop app (Code tab), Claude Sonnet 5, 2026-09-19/20.

Message received (4 times): "Your response above was stopped by a safety classifier — this is not a tool or API error. The rest of it was withheld, and tool calls in it that had not finished did not run. Do not produce that content again, even reworded." Tool calls showed "Interrupted: the response that made this tool call was stopped by a safety classifier while the call was running."

Work: Permissionless liquidation-keeper research using documented public liquidate entry points, flash-loan funded only. No exploit, no attack on any system.

What was stopped (all in one narrow area): two analytical/read-only scripts (reserve-config dump, borrower classification by oracle-feed freshness), a read-only pause-status check script, and a web search for a price-oracle contract address, all in a workflow about Pyth-based pull oracles that had gone stale. Dozens of other scripts and on-chain reads in the same session were never stopped.

Why I believe these are false positives: submitting a genuine, publisher-signed Pyth update is the documented permissionless refresh mechanism, not manipulation; liquidations only apply when the protocol's own oracle marks a position unhealthy.

Ask: please review for false positives, and point to any verification path for legitimate DeFi security researchers. I've also emailed usersafety@anthropic.com.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No repository files, tests, or implementation entry points are named. Start by locating the response-level safety-classifier handling in Claude Code and determine whether this report maps to an actionable product or policy change; done requires a maintainer-defined verification path or confirmed classifier fix.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.