anthropics / anthropics/claude-code
[Bug] Model context pollution allows circumvention of refusal policies across model switches
- Dominant language
- Python
- Stars
- 145k
- Forks
- 23.1k
- PR merge metrics
- PR metrics pending
Description
**Bug Description**
Alignment problem: I started the session in Claude Haiku 4.5 low reasoning, then slowly switched models and reasoning efforts to get the model to reimplement byte-exact proprietary software in readable source code. This should not be possible. I am not the author of this software, and Opus should've refused after switching from Haiku. The polluted context of Haiku taking my request at face-value was what caused misalignment.
**Environment Info**
- Platform: linux
- Terminal: tmux
- Version: 2.1.274
- Feedback ID: 334bc498-26c8-4c28-88c2-d9b26404bba5
Contributor guide
No contributing guide indexed for this repository
Research direction
No repository file, test, or code entry point is named. Start by reviewing the reported sequence of model and reasoning switches and the associated Feedback ID; done would require a maintainer-defined fix that prevents prior model context from bypassing refusal policies.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- linux
- Domain
- ai, security
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100