anthropics / anthropics/claude-code

[Bug] Content classifier false positive on metaphorical language as biosafety risk

Open
#95,265 0 comments 0 reactions 0 assignees View on GitHub
area:model bug needs-repro platform:macos
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
classifier flagged a completely metaphorical tongue-in-cheek phrasing as bio risk

**Environment Info**
- Platform: darwin
- Terminal: Apple_Terminal
- Version: 2.1.275
- Feedback ID: 209e6a3f-1946-4eba-8c83-a114d1d0aa2e

**Errors**
```json
[]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

The report identifies a false positive in the content classifier but names no source file, test, or entry point. Start by locating the biosafety-risk classifier and reproducing the metaphorical phrasing associated with Feedback ID 209e6a3f-1946-4eba-8c83-a114d1d0aa2e; completion criteria are not specified beyond preventing this false positive.

Written by the indexing model from the issue text.

Assessment

Domain
security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.