anthropics / anthropics/claude-code
[cyber] guardrail false positives on legitimate security tooling - auto-reported
- Lenguaje dominante
- Python
- Estrellas
- 145k
- Forks
- 23.1k
- Métricas de merge de PR
- Métricas de PR pendientes
Descripción
## Problem
I am a security researcher building defensive tooling (domain trust auditing, subdomain takeover detection, vulnerability research). The [cyber] guardrail is triggering on every prompt in my workflow, despite the work being entirely legitimate and defensive in nature.
## Impact
- Every interaction forces a model fallback, degrading capability
- Hours of productive work lost to guardrail interruptions
- Feedback submitted through /feedback and support channels with no response or improvement
## What this issue tracks
This issue receives an automated comment for each subsequent false positive occurrence. The comment count demonstrates the volume of the problem. Each comment includes a timestamp, the configured vs actual model, and the session context.
## Context
- Tools being built: domain trust chain analysis, DNS record auditing, mail configuration security checks
- All work is defensive security, identifying vulnerabilities in infrastructure the researcher is authorized to test
- The [cyber] category is too broad, catching standard security research workflows
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Evaluación
Este issue todavía no se ha evaluado.