anthropics / anthropics/claude-code

plugin eval: scope `target: trace` to the graded run's own events

未关闭
#92,017 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area:plugins enhancement
主要语言
Python
星标
145k
派生
23.1k
PR 合并指标
PR 指标待抓取

描述

**Is your feature request related to a problem?**

A `target: trace` regex grader reads the whole trace file. A trace can hold more than one
terminating `result` record: across 125 trace files from our runs we found 23 with more than one,
including one holding subtypes `success`, `error_max_turns`, `success` — all three under a single
session id, in a trace whose own run reported `turns: 6` and `error: null`. A grader asserting
`not_contains: error_(max_turns|during_execution|max_budget_usd)`, which is the only way to detect
an abort from inside a grader, then fails a run that finished normally, with a complete reply.

We have seen this once across roughly 125 runs; it always fails in the safe direction, but it puts
a spurious red on the gate.

**Describe the solution you'd like**

Either scope `target: trace` to the events of the run being graded (the last session, or the
events between the run's own start and its terminating `result`), or expose the run record's
`error` and result `subtype` as a grader target directly, so an abort detector reads one field
instead of grepping a file.

**Describe alternatives you've considered**

Reading the run record's `error` field in a post-run script (we do this for authentication
failures, which end with a `success` subtype and are invisible to any grader target). Telling
readers to check `error` before believing a detector red, which is a note for a human rather than
a fix.

**Additional context**

CLI 2.1.260, `plugin eval` under `CLAUDE_CODE_WALNUT_SPIRE`.

We do not know what causes a single trace file to accumulate several terminating `result` records
under one session id — that may itself be the underlying bug.

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start at the plugin eval `target: trace` grader and reproduce the issue with a trace file containing multiple terminating `result` records. Inspect how the graded run's events and run record are selected; done means the grader can reliably inspect only that run's events or directly target its `error` and result `subtype`.

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
cli, testing-qa
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
活跃
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。