anthropics / anthropics/claude-code

[Bug] Incorrect content moderation classification: legitimate questions marked as cyber

Open
#94,777 0 comments 0 reactions 0 assignees View on GitHub
area:model bug platform:macos
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
Request ID: req_011Cf7E44XaAQE1DwrfZdn4D
Request ID: req_011Cf7DsJMxULNUPA2US82z6
I do automation QA for my company. previously we discussed with Claude questions he has about rules in reports. now it marks questions as cyber for unclear reason

**Environment Info**
- Platform: darwin
- Terminal: iTerm.app
- Version: 2.1.267
- Feedback ID: 2927e0ab-a430-4787-b5cf-bb892ee25086

**Errors**
```json
[]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by investigating the two reported request IDs and Feedback ID 2927e0ab-a430-4787-b5cf-bb892ee25086, using the reported macOS and iTerm environment as context. Reproduce the classification of legitimate report-related questions and determine what behavior distinguishes them from cyber-related requests; done means the false-positive classification is explained and corrected.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.