anthropics / anthropics/claude-code
[cyber] guardrail false positives on legitimate security tooling - auto-reported
- Lingua principale
- Python
- Stelle
- 145k
- Fork
- 23.1k
- Metriche di merge delle PR
- Metriche PR in attesa
Descrizione
## Problem
I am a security researcher building defensive tooling (domain trust auditing, subdomain takeover detection, vulnerability research). The [cyber] guardrail is triggering on every prompt in my workflow, despite the work being entirely legitimate and defensive in nature.
## Impact
- Every interaction forces a model fallback, degrading capability
- Hours of productive work lost to guardrail interruptions
- Feedback submitted through /feedback and support channels with no response or improvement
## What this issue tracks
This issue receives an automated comment for each subsequent false positive occurrence. The comment count demonstrates the volume of the problem. Each comment includes a timestamp, the configured vs actual model, and the session context.
## Context
- Tools being built: domain trust chain analysis, DNS record auditing, mail configuration security checks
- All work is defensive security, identifying vulnerabilities in infrastructure the researcher is authorized to test
- The [cyber] category is too broad, catching standard security research workflows
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.