anthropics / anthropics/claude-code

[cyber] guardrail false positives on legitimate security tooling - auto-reported

オープン
#94,366 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
145k
フォーク
23.1k
PR マージ指標
PR 指標を取得中

説明

## Problem

I am a security researcher building defensive tooling (domain trust auditing, subdomain takeover detection, vulnerability research). The [cyber] guardrail is triggering on every prompt in my workflow, despite the work being entirely legitimate and defensive in nature.

## Impact

- Every interaction forces a model fallback, degrading capability
- Hours of productive work lost to guardrail interruptions
- Feedback submitted through /feedback and support channels with no response or improvement

## What this issue tracks

This issue receives an automated comment for each subsequent false positive occurrence. The comment count demonstrates the volume of the problem. Each comment includes a timestamp, the configured vs actual model, and the session context.

## Context

- Tools being built: domain trust chain analysis, DNS record auditing, mail configuration security checks
- All work is defensive security, identifying vulnerabilities in infrastructure the researcher is authorized to test
- The [cyber] category is too broad, catching standard security research workflows

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。