anthropics / anthropics/claude-code
[Bug] Incorrect content moderation classification: legitimate questions marked as cyber
- Dominant language
- Python
- Stars
- 145k
- Forks
- 23.1k
- PR merge metrics
- PR metrics pending
Description
**Bug Description**
Request ID: req_011Cf7E44XaAQE1DwrfZdn4D
Request ID: req_011Cf7DsJMxULNUPA2US82z6
I do automation QA for my company. previously we discussed with Claude questions he has about rules in reports. now it marks questions as cyber for unclear reason
**Environment Info**
- Platform: darwin
- Terminal: iTerm.app
- Version: 2.1.267
- Feedback ID: 2927e0ab-a430-4787-b5cf-bb892ee25086
**Errors**
```json
[]
```
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by investigating the two reported request IDs and Feedback ID 2927e0ab-a430-4787-b5cf-bb892ee25086, using the reported macOS and iTerm environment as context. Reproduce the classification of legitimate report-related questions and determine what behavior distinguishes them from cyber-related requests; done means the false-positive classification is explained and corrected.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100