anthropics / anthropics/claude-code
[Bug] Custom validator triggers false positive spam detection
- Lenguaje dominante
- Python
- Estrellas
- 145k
- Forks
- 23.1k
- Métricas de merge de PR
- Métricas de PR pendientes
Descripción
**Bug Description**
I was testing a custom anti-spam form validator and while Fable 5.1 was trying to trigger the validation, The model decided I was a bad actor. :-(
Response = "Fable 5's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate coding, cybersecurity, and biology tasks. Switched to Opus 4.8. Send feedback with /feedback or learn more"
Details: `[cyber]`"
**Environment Info**
- Platform: darwin
- Terminal: iTerm.app
- Version: 2.1.252
- Feedback ID: 3ba27065-4d21-4a5d-9c95-9532c8ff546b
**Errors**
```json
[]
```
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
The report names no source file, test, or entry point. Start by reproducing the custom validator flow on darwin in iTerm.app with version 2.1.252, using the supplied feedback ID to trace the safeguard decision; done means the legitimate validation request is no longer flagged as spam.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- ai, security
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Activo
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 35/100