anthropics / anthropics/claude-code
plugin eval: scope `target: trace` to the graded run's own events
- Langage dominant
- Python
- Étoiles
- 145k
- Forks
- 23.1k
- Métriques de merge des PR
- Métriques de PR en attente
Description
**Is your feature request related to a problem?**
A `target: trace` regex grader reads the whole trace file. A trace can hold more than one
terminating `result` record: across 125 trace files from our runs we found 23 with more than one,
including one holding subtypes `success`, `error_max_turns`, `success` — all three under a single
session id, in a trace whose own run reported `turns: 6` and `error: null`. A grader asserting
`not_contains: error_(max_turns|during_execution|max_budget_usd)`, which is the only way to detect
an abort from inside a grader, then fails a run that finished normally, with a complete reply.
We have seen this once across roughly 125 runs; it always fails in the safe direction, but it puts
a spurious red on the gate.
**Describe the solution you'd like**
Either scope `target: trace` to the events of the run being graded (the last session, or the
events between the run's own start and its terminating `result`), or expose the run record's
`error` and result `subtype` as a grader target directly, so an abort detector reads one field
instead of grepping a file.
**Describe alternatives you've considered**
Reading the run record's `error` field in a post-run script (we do this for authentication
failures, which end with a `success` subtype and are invisible to any grader target). Telling
readers to check `error` before believing a detector red, which is a note for a human rather than
a fix.
**Additional context**
CLI 2.1.260, `plugin eval` under `CLAUDE_CODE_WALNUT_SPIRE`.
We do not know what causes a single trace file to accumulate several terminating `result` records
under one session id — that may itself be the underlying bug.
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Piste de recherche
Start at the plugin eval `target: trace` grader and reproduce the issue with a trace file containing multiple terminating `result` records. Inspect how the graded run's events and run record are selected; done means the grader can reliably inspect only that run's events or directly target its `error` and result `subtype`.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- cli, testing-qa
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Active
- Clarté
- Plutôt claire
- Accessibilité débutants
- 35/100