anthropics / anthropics/claude-code

[Bug] Sonnet 5 [bio] classifier false positive on Claude.ai web chat — DHCP/DNS/AD infrastructure content (NOT Claude Code CLI)

Open
#95,602 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

invalid
Dominant language
TypeScript
Stars
147k
Forks
24k
PR merge metrics
PR metrics pending

Description

IMPORTANT: This occurred on Claude.ai (web chat), not the Claude Code CLI.
I'm filing here because Anthropic's own documentation indicates the safety
classifier is shared across products (Claude.ai, Cowork, Claude Code, and
the API each have separately reported fallback rates for this same
classifier). If this is the wrong venue for a Claude.ai-specific report,
please redirect me to the correct one — the support.anthropic.com portal
and support@anthropic.com have both failed to produce a response after
7+ days.

Preflight Checklist
  • I have searched existing issues and this hasn't been reported yet
  • This is a single bug report (please file separate reports for different bugs)
  • I am using the latest version of Claude Code
What's Wrong?

Bug Description

Sonnet 5, via Claude.ai (web chat, not Claude Code CLI), hard-blocked a legitimate
IT infrastructure documentation project with a classifier tag of "BIO." The content
was a Windows Server / Active Directory migration runbook covering DHCP scope
migration and static DNS client remediation — standard enterprise IT vocabulary,
zero biological content.

Environment

  • Product: Claude.ai (web chat), not Claude Code CLI
  • Model: Claude Sonnet 5
  • Platform: [Firefox Windows 10 Enterprise LTSC]
  • 9-12
  • Feedback ID shown on the blocked message: [fill in if one was displayed, otherwise write "none shown"]

What Was Flagged

The blocked content was a Track B execution runbook (DHCP transition, static DNS
client remediation, endpoint security planning) for an Active Directory 2022
upgrade project. Suspect trigger vocabulary, all used in the ordinary IT sense:
"remediation," "host," "scope," "relay," "sweep," and — in a follow-up attempt
to discuss endpoint security — "malware," "infection," "propagation."

Why This Is a False Positive

No biological, chemical, or weapons content of any kind appears anywhere in the
source material. All terms are standard Windows Server/networking terminology.
The identical subject matter is handled without any block on Claude Sonnet 4.6
in the same account/session context.

Additional Impact — Sticky Block / No Recourse

  • The block was an INPUT-level classifier: no response was generated at all, and
    no thumbs-down/feedback control was available in the UI (distinct from an
    output-classifier flag, which does show a thumbs-down).
  • Once tripped, the topic became unworkable on Sonnet 5 for the rest of the
    engagement — matches the "sticky for the rest of session" pattern described in
    issues #94220 and #94223.
  • Attempted to report through official channels with no resolution:
    • support.anthropic.com: login loop prevents ticket submission (logged in,
      but the "log in" button on the support widget opens a new Claude chat
      instead of authenticating to support)
    • support@anthropic.com: received only an automated Fin bot acknowledgment
      with a queue conversation ID; no human response after 7+ days

Business Impact

Paying business account (company-funded seat), active client-facing
infrastructure consulting engagement blocked mid-project.

Related Issues (same bug class)

#94220, #94223 — Sonnet 5, [bio], sticky block, physics content, no bio subject matter
#95247 — [bio], PostgreSQL database schema, RFP tool
#75668 — [bio], CNC/HMI session, sticky, suspected trigger: manufacturing/dental vocabulary
#95265, #95479 — same pattern, [bio]/[cyber], other benign technical content

This report adds a new, previously unreported vocabulary cluster (DHCP/DNS/AD
networking + endpoint security terminology) to the existing pattern.

Request

  1. Investigate and tune the Sonnet 5 input classifier against enterprise IT/
    networking vocabulary.
  2. Separately: the support.anthropic.com login loop is itself a bug preventing
    ticket submission for logged-in users, and should be triaged independently.
What Should Happen?

I should be able to complete my AD, DHCP and DNS migration runbook

Error Messages/Logs
- Product: Claude.ai (web chat), not Claude Code CLI
- Model: Claude Sonnet 5
- Platform: [Firefox Windows 10]
- Fin Conversation ID for your records: 215475920941428
- Exact error text shown: "Sonnet 5's safeguards flagged this message. Our 
  intentionally broad safeguards allow us to deliver more capabilities faster, 
  but can sometimes flag legitimate coding, cybersecurity, and biology tasks." 
  Details: BIO
- Feedback ID: none shown for this block (this was an input-classifier hard 
  halt, not an output-flagged message with a thumbs-down/feedback option)
Steps to Reproduce

It will halt if I ask it to create the final draft of my migration runbook

Claude Model

Sonnet (default)

Is this a regression?

Yes, this worked in a previous version

Last Working Version

No response

Claude Code Version

N/A

Platform

Other

Operating System

Windows 10 Enterprise

Terminal/Shell

PowerShell

Additional Information

all my work is done in a browser at claude.ai

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No repository files, tests, or implementation entry points are identified; the report concerns Claude.ai rather than the Claude Code CLI. Start by reviewing related issues #94220, #94223, and #95247 and reproducing the block with the described AD, DHCP, and DNS runbook. Done would require the legitimate runbook to proceed without the BIO block and the separate support-login problem to be triaged independently.

Written by the indexing model from the issue text.

Assessment

Tech stack
powershell
Domain
networking, security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.