Follow-up: replay fixtures and evaluation runner for measured route recommendations (#151 AC6-AC8)
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 38/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- ruby
- Domain
- documentation, testing-qa, tooling
Research direction
Start by reading the fixture conventions under skills/post-merge-audit/fixtures/ and test/fixtures/, then inspect skills/pr-batch/bin/model-routing-contract-test.rb and the execution-provenance dependency. Scope the first implementation to the bounded-helper scenario class, with isolated randomized repetitions and raw records. Done means the runner enforces observed receipts and records the listed quality, cost, churn, timing, and route-adherence metrics, with tests and measured results replacing UNKNOWN in the evidence table.
Written by the indexing model from the issue text.
Description
Problem
#151 acceptance criteria 6, 7, and 8 are not closed:
- Deterministic replay fixtures cover #114/#116/#123.
- The evaluation runner supports isolated, randomized, repeatable candidate runs.
- Results include quality, churn, time, tokens/cost, and route-adherence metrics.
docs/agent-workflows-model-routing.md now publishes the scenario taxonomy with every class marked sample count 0 and evidence strength UNKNOWN, and states that both conservative profiles are priors chosen for fail-closed safety rather than measurements. That is the honest current state: no route recommendation in this repo is backed by a measured comparison. Closing that gap needs fixtures and a runner, which are code, not prose.
Until then the profiles stay provisional — claude-profile v0 is explicitly marked as such and pinned across surfaces by skills/pr-batch/bin/model-routing-contract-test.rb, and Fable 5 stays experimental pending this evidence.
Files
- new: replay fixtures for the motivating incidents, following the existing fixture conventions under
skills/post-merge-audit/fixtures/andtest/fixtures/ - new: an evaluation runner plus its unit test under
skills/pr-batch/bin/ docs/agent-workflows-model-routing.md— the "Evidence Status" table is where measured results replaceUNKNOWN, with sample count, date, host version, and model availability- depends on the execution-provenance schema follow-up (route-adherence metrics come from validated receipts)
Direction
Fixtures. Use the incidents #151 names, each with deterministic acceptance tests:
- #114 and PR #31 — current-head pending-review-draft readiness and incomplete-inventory controls
- #116 — dispatcher identity, fallback authorization, persistence, and replacement-fencing state machine
- #123 — prompt headroom, route-group splitting, configured-reviewer gates, and compact/canonical completion-contract mutations
Each run starts from the same base, uses the same sanitized task prompt, runs in an isolated worktree, and receives no findings from another candidate before completion.
Runner. Randomized candidate order, at least 3 independent repetitions per tuple when cost permits, one fixed blinded review route for comparison, and a separate high-versus-xhigh review-effort comparison on identical candidate diffs. Never compare a requested route that lacks an observed receipt — that rule is already normative in the routing doc and the runner must enforce it rather than restate it.
Metrics. Acceptance tests passing before review; P1/P2/P3 findings from blinded review; escaped defects; correction waves and post-publication commits; changed-line churn and reverted work; wall-clock and active-agent time; tokens and estimated cost when available; route-adherence failures; exact-head QA result; maintainer intervention count. Preserve raw run records so recommendations can be recalculated as pricing or models change.
Note the cost shape before starting: this is the most expensive item in #151 by a wide margin, and it is worth scoping to one scenario class end-to-end first to prove the harness, rather than building all seven at once. The bounded-helper class is the cheapest place to validate the mechanism.
Origin
Follow-up from the aw-g batch lane for #151, which owned only docs/agent-workflows-model-routing.md and skills/pr-batch/bin/model-routing-contract-test.rb. Fixtures and a runner were outside those paths, so the lane published the taxonomy with explicit UNKNOWN evidence strength instead of implying measurements that do not exist.
Refs #151
- Dominant language
- Ruby
- Stars
- 7
- Forks
- 1
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 150
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from shakacode/agent-workflows
-
complexity:neutral follow-up P3 triage:park
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
shakacode/agent-workflows#751 ·
-
complexity:complexify follow-up needs-customer-feedback P3 review-nit triage:park
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
shakacode/agent-workflows#750 ·
-
complexity:neutral follow-up P2 triage:reconcile
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
shakacode/agent-workflows#640 ·
-
complexity:neutral P2 triage:reconcile
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
shakacode/agent-workflows#483 · 1 comment ·
-
complexity:complexify enhancement P2
Difficulty 5/5 Over a week Newbie friendliness 25/100
shakacode/agent-workflows#852 ·
All issues in shakacode/agent-workflows
Similar issues
-
バグ
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
voxpupuli/puppet-epel#186 · 1 comment ·
-
external_created_at is no longer used for the message timestamp since the new message UI (v4.4.0) OpenBug Frontend
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
TheOdinProject/curriculum#31402 · 1 comment ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100