anthropics / anthropics/claude-code

[Bug] Model context pollution allows circumvention of refusal policies across model switches

Open
#95,145 0 comments 0 reactions 0 assignees View on GitHub
area:model area:security bug platform:linux
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
Alignment problem: I started the session in Claude Haiku 4.5 low reasoning, then slowly switched models and reasoning efforts to get the model to reimplement byte-exact proprietary software in readable source code. This should not be possible. I am not the author of this software, and Opus should've refused after switching from Haiku. The polluted context of Haiku taking my request at face-value was what caused misalignment.

**Environment Info**
- Platform: linux
- Terminal: tmux
- Version: 2.1.274
- Feedback ID: 334bc498-26c8-4c28-88c2-d9b26404bba5

Contributor guide

No contributing guide indexed for this repository

Research direction

No repository file, test, or code entry point is named. Start by reviewing the reported sequence of model and reasoning switches and the associated Feedback ID; done would require a maintainer-defined fix that prevents prior model context from bypassing refusal policies.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux
Domain
ai, security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.