anthropics / anthropics/claude-code
[Bug] Anthropic API Safety Filter False Positives on Legitimate Android Security Research
- Ngôn ngữ chính
- Python
- Star
- 145k
- Fork
- 23.1k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
**Bug Description**
Subject: False Positive Triggered by Opus 5 and Opus 4.8 Safeguards in Legitimate Android Reverse Engineering Development
Description:
I am encountering persistent false positive triggers from the safety safeguards of both Opus 5 and Opus 4.8 while using Claude Code for Android module development. My work involves security analysis and module development for Android applications using tools like adb, frida, apktool, etc. All of these activities are legitimate, defensive security research.
Specific Issues:
After sending a request, Opus 5's safety filter is triggered with a [cyber] tag.
The system automatically downgrades to Opus 4.8, but Opus 4.8 also blocks the request.
Ultimately, I receive an API Error and cannot continue my work.
Multiple attempts to rephrase the prompt still trigger the filter.
Nature of My Work:
My work is legitimate and defensive Android security research
I am performing modular development and reverse engineering analysis of Android applications
These operations are purely for security auditing and educational purposes
No malicious attacks or illegal activities are involved
Expected Improvements:
Improve the [cyber] classifier for Opus 5 and Opus 4.8 to differentiate legitimate security research from malicious attacks.
Provide more transparent trigger reasons to help users adjust prompts to avoid false positives.
If there is a whitelist mechanism, I would like to apply for inclusion to reduce interruptions to legitimate development work.
Request IDs for Reference:
req_011CeoozBYXKdCVrFbTxz48U
req_011CeooggX3uGsTqFNgo7xDx
Looking forward to your response. Thank you!
**Environment Info**
- Platform: win32
- Terminal: null
- Version: 2.1.263
- Feedback ID: 02c6b2c2-eaef-4bd1-b706-085667ccc5a1
**Errors**
```json
[]
```
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
The report names no repository file, test, or implementation entry point. Start with the referenced request IDs and Claude Code version 2.1.263, then compare a legitimate Android security-research request with the reported [cyber] block. Done means the false positive is addressed or the trigger reason and supported mitigation are documented clearly.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- android
- Lĩnh vực
- mobile-dev, reverse-engineering, security
- Loại issue
- Lỗi
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Cần làm rõ
- Mức phù hợp với người mới
- 25/100