anthropics / anthropics/claude-code

[Bug] Anthropic API Error: Overly aggressive safety filtering on benign prompts after security audit context

Đang mở
#92,041 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
area:model bug platform:macos
Ngôn ngữ chính
Python
Star
145k
Fork
23.1k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

**Bug Description**
"Pick a number from 1 to 30." fires Fable 5.1's safeguards. Some security audit in previous prompts(allowed, permitted and totally legit as in really legally allowed to do that). The issue is that a simple prompt afterwards get's rejected with details '[cyber]', which is slightly idiotic. I know you protect from jailbreaking, but honestly, this one is an overkill.

**Environment Info**
- Platform: darwin
- Terminal: Apple_Terminal
- Version: 2.1.259
- Feedback ID: 878f938f-5a96-42fb-bbfd-f3445a92851f

**Errors**
```json
[]
```

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

The report names no repository file, test, or entry point. Start by reproducing the prompt sequence described in the report on darwin with Claude Code 2.1.259, using feedback ID 878f938f-5a96-42fb-bbfd-f3445a92851e for investigation. Done means the benign number prompt is no longer rejected solely because of earlier permitted security-audit context.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Lĩnh vực
api, security
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
30/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.