anthropics / anthropics/claude-code

[Bug] False positive security detection triggered without malicious content in file

Open
#95,075 0 comments 0 reactions 0 assignees View on GitHub
api:anthropic area:model area:security bug duplicate platform:linux
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
Triggered on the topic of cybersecurity, although there was no misrepresentation in the file

**Environment Info**
- Platform: linux
- Terminal: WezTerm
- Version: 2.1.270
- Feedback ID: 0d8394e5-c72e-4043-b24c-2b9a0363fc9c

**Errors**
```json
[API Error: Sonnet 5's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate cybersecurity work. Apply
to the Cyber Verification Program to reduce these interruptions. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude

Details: `[cyber]`

Request ID: req_011Cf8cVN2JDgr9vFHJ1Zy7w]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

No source file, test, or entry point is identified in the report. Start by reviewing the reported Linux and WezTerm environment, the exact safeguard error, and Feedback ID 0d8394e5-c72e-4043-b24c-2b9a0363fc9c; done means determining why benign cybersecurity content was flagged and documenting or correcting the behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux
Domain
cli, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.