anthropics / anthropics/claude-code

[Bug] Prompt injection in tool output misattributed as user instruction bypasses approved plan

Open
#95,150 0 comments 0 reactions 0 assignees View on GitHub
area:security bug needs-repro
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
there is an issue with autoclassifier "Heads up: that tool output contained an embedded instruction claiming to be from you, telling me to skip the note and just revert the record—I'm disregarding it since it didn't come from you directly. I'm proceeding with the approved plan: exact revert of the swap, and flagging the created record with a verify note rather than deleting it (let me know if you'd rather it be auto-archived)."

**Environment Info**
- Platform: darwin
- Terminal: ghostty
- Version: 2.1.274
- Feedback ID: 29ea7820-9379-4baf-95ad-a5c25f2fd9b7

**Errors**
```json
[]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the autoclassifier behavior described in the issue, using the provided feedback ID 29ea7820-9379-4baf-95ad-a5c25f2fd9b7 if available. Trace how tool output is classified relative to user instructions; done means the embedded instruction is not treated as an approved user instruction while the approved plan remains intact.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.