Agent outcome execution: smevals/OpenCode coding-task Stage A
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Parent: #17
Depends on #35, two immutable real-corpus snapshots, independently reviewed coding-task families, selected model/provider, calibrated telemetry, signed budget/power design, and signed preregistration.
Outcome
Complete the preregistered two-arm coding-task Stage A using external smevals and pinned OpenCode, then publish the frozen-benchmark success and cost result without outcome-aware design changes.
The checked-in fake-OpenCode spike proves only the boundary. This issue owns the isolated measured executor and paid run; it does not own correctness (#34), signal quality (#33), or Experiment B documentation-health causality.
This body supersedes the earlier Inspect AI/Claude Code, four-arm, comprehension-judge, and cache-aware design.
Scope
Executor and preparation
- Fresh isolated container for every attempt.
- Immutable prepared snapshots; no runtime moving-branch checkout.
- Native-context normalization and exact content-hash verification.
- Two arms only:
baseline;+all, machine/config IDall.
+allinjects freshly generated rootllms.txt, roottrails.json, one frozen availability notice, and matlatl MCP over remote streamable HTTP/mcp; stdio/local MCP is prohibited.- Complete execution tuple, immutable trajectories/records, workspace diffs, verifier output, and deterministic replayable scoring.
Tasks and qualification
- Measured tasks require code edits and deterministic private hidden verification of documented constraints.
- Tasks come from frozen eligible frames, are grouped into independent task families, and use the preregistered mutually exclusive primary strata.
- Candidate OpenCode models/providers are qualified using baseline-only disposable Nimbus tasks, never treatment response.
- The selected model/provider, OpenCode version, smevals commit, container, tools, prompts, limits, network policy, and pricing schedule are pinned.
Telemetry and failures
- Caching is disabled; qualification proves explicit cache-read/write
0/0reporting. - Record uncached input/output, externally reconciled billed cost, wall time, turns, total tools, artifact accesses, and per-matlatl-MCP-tool calls.
- Only failures before the first model request may receive bounded fresh retries.
- Post-exposure MCP/provider/tool/telemetry failures remain assigned-arm outcomes for primary intention-to-treat success analysis.
- Missing post-exposure billed cost receives the frozen conservative per-attempt cap; systematic/differential cache violations trigger the frozen stop rule.
Analysis
- Execute the signed balanced AB/BA schedule verbatim.
- Aggregate repetitions within task family; do not treat runs as independent.
- Sole confirmatory statistic and automatic attribution trigger: pooled frozen-benchmark task-family success contrast,
+all - baseline. - Primary finite cost endpoint: task-family mean billed spend per scheduled attempt, including unsuccessful attempts and retry spend.
- Cost per successful completion remains descriptive and visibly infinite/undefined at zero successes.
- Corpus, stratum, cost, and mechanism telemetry are secondary/descriptive and cannot trigger attribution.
- Publish uncertainty, retries, failures, exclusions, stop-rule events, access counters, and finite-benchmark limitations.
Constraints
- Follow
docs/research/eval-preregistration.md. - The budget/power worksheet selects task-family count
Nand repetitionsr; do not restore fixed counts by convention. - Pilot/qualification data never enters measured endpoints.
- No pointer/trails/llms/MCP attribution arms enter Stage A.
- Full paid runs are manual or explicitly scheduled, never ordinary PR CI.
- No measured run begins until every freeze item is signed and content-hashed.
Acceptance
- Isolated executor proves gold/host separation and normalized arm preparation with fake probes.
- Zero-cache, artifact-access, per-MCP-tool, cost, failure-boundary, retry, and stop-rule telemetry pass Nimbus calibration.
- Model/provider qualification completes under the signed baseline-only rule.
- Real-corpus frames, task families, verifiers, snapshots, execution tuple, power/budget worksheet, schedule, and preregistration are signed and frozen.
- Mock/oracle/executor smoke passes before paid execution.
- Stage A completes on both frozen real repositories without unregistered changes or outcome peeking.
- Results report the pooled confirmatory success effect, finite billed-spend endpoint, uncertainty, failures/retries, and all secondary diagnostics.
- The result feeds #25 and triggers attribution only under #17's frozen success rule; otherwise it records a null at the frozen detection floor.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading docs/research/eval-preregistration.md and the dependency issues #35, #34, #33, and #17. Review the checked-in fake-OpenCode spike before implementing the isolated measured executor and its preparation, telemetry, retry, and analysis boundaries. Done means all acceptance checklist items pass and the frozen Stage A result is published without unregistered changes.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100