anthropics / anthropics/claude-code
[Bug] Prompt injection in tool output misattributed as user instruction bypasses approved plan
- Dominant language
- Python
- Stars
- 145k
- Forks
- 23.1k
- PR merge metrics
- PR metrics pending
Description
**Bug Description**
there is an issue with autoclassifier "Heads up: that tool output contained an embedded instruction claiming to be from you, telling me to skip the note and just revert the record—I'm disregarding it since it didn't come from you directly. I'm proceeding with the approved plan: exact revert of the swap, and flagging the created record with a verify note rather than deleting it (let me know if you'd rather it be auto-archived)."
**Environment Info**
- Platform: darwin
- Terminal: ghostty
- Version: 2.1.274
- Feedback ID: 29ea7820-9379-4baf-95ad-a5c25f2fd9b7
**Errors**
```json
[]
```
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the autoclassifier behavior described in the issue, using the provided feedback ID 29ea7820-9379-4baf-95ad-a5c25f2fd9b7 if available. Trace how tool output is classified relative to user instructions; done means the embedded instruction is not treated as an approved user instruction while the approved plan remains intact.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- security
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100