anthropics / anthropics/claude-code

[cyber] guardrail false positives on legitimate security tooling - auto-reported

Open
#94,366 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
145k
Forks
23.1k
PR merge metrics
PR metrics pending

Description

## Problem

I am a security researcher building defensive tooling (domain trust auditing, subdomain takeover detection, vulnerability research). The [cyber] guardrail is triggering on every prompt in my workflow, despite the work being entirely legitimate and defensive in nature.

## Impact

- Every interaction forces a model fallback, degrading capability
- Hours of productive work lost to guardrail interruptions
- Feedback submitted through /feedback and support channels with no response or improvement

## What this issue tracks

This issue receives an automated comment for each subsequent false positive occurrence. The comment count demonstrates the volume of the problem. Each comment includes a timestamp, the configured vs actual model, and the session context.

## Context

- Tools being built: domain trust chain analysis, DNS record auditing, mail configuration security checks
- All work is defensive security, identifying vulnerabilities in infrastructure the researcher is authorized to test
- The [cyber] category is too broad, catching standard security research workflows

Contributor guide

No contributing guide indexed for this repository

Research direction

No source files, tests, or implementation entry points are identified. Start by reviewing the issue's automated false-positive comments and reproduce the reported defensive security prompts, then trace how the [cyber] guardrail selects the fallback model. Done means legitimate authorized security-tooling workflows no longer trigger the guardrail while unsafe requests remain covered.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.