anthropics / anthropics/claude-code

plugin eval: scope `target: trace` to the graded run's own events

オープン
#92,017 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
area:plugins enhancement
主要言語
Python
スター
145k
フォーク
23.1k
PR マージ指標
PR 指標を取得中

説明

**Is your feature request related to a problem?**

A `target: trace` regex grader reads the whole trace file. A trace can hold more than one
terminating `result` record: across 125 trace files from our runs we found 23 with more than one,
including one holding subtypes `success`, `error_max_turns`, `success` — all three under a single
session id, in a trace whose own run reported `turns: 6` and `error: null`. A grader asserting
`not_contains: error_(max_turns|during_execution|max_budget_usd)`, which is the only way to detect
an abort from inside a grader, then fails a run that finished normally, with a complete reply.

We have seen this once across roughly 125 runs; it always fails in the safe direction, but it puts
a spurious red on the gate.

**Describe the solution you'd like**

Either scope `target: trace` to the events of the run being graded (the last session, or the
events between the run's own start and its terminating `result`), or expose the run record's
`error` and result `subtype` as a grader target directly, so an abort detector reads one field
instead of grepping a file.

**Describe alternatives you've considered**

Reading the run record's `error` field in a post-run script (we do this for authentication
failures, which end with a `success` subtype and are invisible to any grader target). Telling
readers to check `error` before believing a detector red, which is a note for a human rather than
a fix.

**Additional context**

CLI 2.1.260, `plugin eval` under `CLAUDE_CODE_WALNUT_SPIRE`.

We do not know what causes a single trace file to accumulate several terminating `result` records
under one session id — that may itself be the underlying bug.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

Start at the plugin eval `target: trace` grader and reproduce the issue with a trace file containing multiple terminating `result` records. Inspect how the graded run's events and run record are selected; done means the grader can reliably inspect only that run's events or directly target its `error` and result `subtype`.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
cli, testing-qa
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。