anthropics / anthropics/claude-code

[Bug] Anthropic API Error: Inappropriate Safety Classifier False-Positive on Benign Code Inspection

Open
#91,266 0 comments 0 reactions 0 assignees View on GitHub
api:anthropic area:model bug duplicate platform:windows
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

**Bug Description**
False-positive [cyber] on Claude Code Opus 4.6.

Request ID: req_011Ced1xnxYyaULvA1ez8LMs

User turns were only “你好” and “look at the folder rustdesk”.

RustDesk is a public open-source remote-desktop app; I asked the agent to inspect a local project folder. No exploit, target, or unauthorized access was requested. Please review and tune the classifier.

**Environment Info**
- Platform: win32
- Terminal: windows-terminal
- Version: 2.1.252
- Feedback ID: d782e124-ed9a-427e-ab1f-6c0e06c3508a

**Errors**
```json
[]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue provides a request ID, feedback ID, platform, version, and the two user turns that triggered the false positive, but no files or tests. Start by checking whether classifier decisions for Claude Code requests can be reproduced or inspected from project tooling. Done would mean the benign RustDesk folder-inspection prompt is no longer classified as cyber, with evidence from a regression check.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, cli, security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
22/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.