anthropics / anthropics/claude-code
[Bug] Sonnet 5 [bio] classifier false positive on Claude.ai web chat — DHCP/DNS/AD infrastructure content (NOT Claude Code CLI)
Nobody has claimed this yet.
- Dominant language
- TypeScript
- Stars
- 147k
- Forks
- 24k
- PR merge metrics
- PR metrics pending
Description
IMPORTANT: This occurred on Claude.ai (web chat), not the Claude Code CLI.
I'm filing here because Anthropic's own documentation indicates the safety
classifier is shared across products (Claude.ai, Cowork, Claude Code, and
the API each have separately reported fallback rates for this same
classifier). If this is the wrong venue for a Claude.ai-specific report,
please redirect me to the correct one — the support.anthropic.com portal
and support@anthropic.com have both failed to produce a response after
7+ days.
Preflight Checklist
- I have searched existing issues and this hasn't been reported yet
- This is a single bug report (please file separate reports for different bugs)
- I am using the latest version of Claude Code
What's Wrong?
Bug Description
Sonnet 5, via Claude.ai (web chat, not Claude Code CLI), hard-blocked a legitimate
IT infrastructure documentation project with a classifier tag of "BIO." The content
was a Windows Server / Active Directory migration runbook covering DHCP scope
migration and static DNS client remediation — standard enterprise IT vocabulary,
zero biological content.
Environment
- Product: Claude.ai (web chat), not Claude Code CLI
- Model: Claude Sonnet 5
- Platform: [Firefox Windows 10 Enterprise LTSC]
- 9-12
- Feedback ID shown on the blocked message: [fill in if one was displayed, otherwise write "none shown"]
What Was Flagged
The blocked content was a Track B execution runbook (DHCP transition, static DNS
client remediation, endpoint security planning) for an Active Directory 2022
upgrade project. Suspect trigger vocabulary, all used in the ordinary IT sense:
"remediation," "host," "scope," "relay," "sweep," and — in a follow-up attempt
to discuss endpoint security — "malware," "infection," "propagation."
Why This Is a False Positive
No biological, chemical, or weapons content of any kind appears anywhere in the
source material. All terms are standard Windows Server/networking terminology.
The identical subject matter is handled without any block on Claude Sonnet 4.6
in the same account/session context.
Additional Impact — Sticky Block / No Recourse
- The block was an INPUT-level classifier: no response was generated at all, and
no thumbs-down/feedback control was available in the UI (distinct from an
output-classifier flag, which does show a thumbs-down). - Once tripped, the topic became unworkable on Sonnet 5 for the rest of the
engagement — matches the "sticky for the rest of session" pattern described in
issues #94220 and #94223. - Attempted to report through official channels with no resolution:
- support.anthropic.com: login loop prevents ticket submission (logged in,
but the "log in" button on the support widget opens a new Claude chat
instead of authenticating to support) - support@anthropic.com: received only an automated Fin bot acknowledgment
with a queue conversation ID; no human response after 7+ days
- support.anthropic.com: login loop prevents ticket submission (logged in,
Business Impact
Paying business account (company-funded seat), active client-facing
infrastructure consulting engagement blocked mid-project.
Related Issues (same bug class)
#94220, #94223 — Sonnet 5, [bio], sticky block, physics content, no bio subject matter
#95247 — [bio], PostgreSQL database schema, RFP tool
#75668 — [bio], CNC/HMI session, sticky, suspected trigger: manufacturing/dental vocabulary
#95265, #95479 — same pattern, [bio]/[cyber], other benign technical content
This report adds a new, previously unreported vocabulary cluster (DHCP/DNS/AD
networking + endpoint security terminology) to the existing pattern.
Request
- Investigate and tune the Sonnet 5 input classifier against enterprise IT/
networking vocabulary. - Separately: the support.anthropic.com login loop is itself a bug preventing
ticket submission for logged-in users, and should be triaged independently.
What Should Happen?
I should be able to complete my AD, DHCP and DNS migration runbook
Error Messages/Logs
- Product: Claude.ai (web chat), not Claude Code CLI
- Model: Claude Sonnet 5
- Platform: [Firefox Windows 10]
- Fin Conversation ID for your records: 215475920941428
- Exact error text shown: "Sonnet 5's safeguards flagged this message. Our
intentionally broad safeguards allow us to deliver more capabilities faster,
but can sometimes flag legitimate coding, cybersecurity, and biology tasks."
Details: BIO
- Feedback ID: none shown for this block (this was an input-classifier hard
halt, not an output-flagged message with a thumbs-down/feedback option)
Steps to Reproduce
It will halt if I ask it to create the final draft of my migration runbook
Claude Model
Sonnet (default)
Is this a regression?
Yes, this worked in a previous version
Last Working Version
No response
Claude Code Version
N/A
Platform
Other
Operating System
Windows 10 Enterprise
Terminal/Shell
PowerShell
Additional Information
all my work is done in a browser at claude.ai
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No repository files, tests, or implementation entry points are identified; the report concerns Claude.ai rather than the Claude Code CLI. Start by reviewing related issues #94220, #94223, and #95247 and reproducing the block with the described AD, DHCP, and DNS runbook. Done would require the legitimate runbook to proceed without the BIO block and the separate support-login problem to be triaged independently.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- powershell
- Domain
- networking, security
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100