anthropics / anthropics/claude-code

[cyber] guardrail false positives on legitimate security tooling - auto-reported

Abierto
#94,366 1 comentario 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
145k
Forks
23.1k
Métricas de merge de PR
Métricas de PR pendientes

Descripción

## Problem

I am a security researcher building defensive tooling (domain trust auditing, subdomain takeover detection, vulnerability research). The [cyber] guardrail is triggering on every prompt in my workflow, despite the work being entirely legitimate and defensive in nature.

## Impact

- Every interaction forces a model fallback, degrading capability
- Hours of productive work lost to guardrail interruptions
- Feedback submitted through /feedback and support channels with no response or improvement

## What this issue tracks

This issue receives an automated comment for each subsequent false positive occurrence. The comment count demonstrates the volume of the problem. Each comment includes a timestamp, the configured vs actual model, and the session context.

## Context

- Tools being built: domain trust chain analysis, DNS record auditing, mail configuration security checks
- All work is defensive security, identifying vulnerabilities in infrastructure the researcher is authorized to test
- The [cyber] category is too broad, catching standard security research workflows

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.