stacklok / stacklok/matlatl

Agent outcome execution: smevals/OpenCode coding-task Stage A

Open
#36 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Go
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Parent: #17

Depends on #35, two immutable real-corpus snapshots, independently reviewed coding-task families, selected model/provider, calibrated telemetry, signed budget/power design, and signed preregistration.

Outcome

Complete the preregistered two-arm coding-task Stage A using external smevals and pinned OpenCode, then publish the frozen-benchmark success and cost result without outcome-aware design changes.

The checked-in fake-OpenCode spike proves only the boundary. This issue owns the isolated measured executor and paid run; it does not own correctness (#34), signal quality (#33), or Experiment B documentation-health causality.

This body supersedes the earlier Inspect AI/Claude Code, four-arm, comprehension-judge, and cache-aware design.

Scope

Executor and preparation
  • Fresh isolated container for every attempt.
  • Immutable prepared snapshots; no runtime moving-branch checkout.
  • Native-context normalization and exact content-hash verification.
  • Two arms only:
    • baseline;
    • +all, machine/config ID all.
  • +all injects freshly generated root llms.txt, root trails.json, one frozen availability notice, and matlatl MCP over remote streamable HTTP /mcp; stdio/local MCP is prohibited.
  • Complete execution tuple, immutable trajectories/records, workspace diffs, verifier output, and deterministic replayable scoring.
Tasks and qualification
  • Measured tasks require code edits and deterministic private hidden verification of documented constraints.
  • Tasks come from frozen eligible frames, are grouped into independent task families, and use the preregistered mutually exclusive primary strata.
  • Candidate OpenCode models/providers are qualified using baseline-only disposable Nimbus tasks, never treatment response.
  • The selected model/provider, OpenCode version, smevals commit, container, tools, prompts, limits, network policy, and pricing schedule are pinned.
Telemetry and failures
  • Caching is disabled; qualification proves explicit cache-read/write 0/0 reporting.
  • Record uncached input/output, externally reconciled billed cost, wall time, turns, total tools, artifact accesses, and per-matlatl-MCP-tool calls.
  • Only failures before the first model request may receive bounded fresh retries.
  • Post-exposure MCP/provider/tool/telemetry failures remain assigned-arm outcomes for primary intention-to-treat success analysis.
  • Missing post-exposure billed cost receives the frozen conservative per-attempt cap; systematic/differential cache violations trigger the frozen stop rule.
Analysis
  • Execute the signed balanced AB/BA schedule verbatim.
  • Aggregate repetitions within task family; do not treat runs as independent.
  • Sole confirmatory statistic and automatic attribution trigger: pooled frozen-benchmark task-family success contrast, +all - baseline.
  • Primary finite cost endpoint: task-family mean billed spend per scheduled attempt, including unsuccessful attempts and retry spend.
  • Cost per successful completion remains descriptive and visibly infinite/undefined at zero successes.
  • Corpus, stratum, cost, and mechanism telemetry are secondary/descriptive and cannot trigger attribution.
  • Publish uncertainty, retries, failures, exclusions, stop-rule events, access counters, and finite-benchmark limitations.

Constraints

  • Follow docs/research/eval-preregistration.md.
  • The budget/power worksheet selects task-family count N and repetitions r; do not restore fixed counts by convention.
  • Pilot/qualification data never enters measured endpoints.
  • No pointer/trails/llms/MCP attribution arms enter Stage A.
  • Full paid runs are manual or explicitly scheduled, never ordinary PR CI.
  • No measured run begins until every freeze item is signed and content-hashed.

Acceptance

  • Isolated executor proves gold/host separation and normalized arm preparation with fake probes.
  • Zero-cache, artifact-access, per-MCP-tool, cost, failure-boundary, retry, and stop-rule telemetry pass Nimbus calibration.
  • Model/provider qualification completes under the signed baseline-only rule.
  • Real-corpus frames, task families, verifiers, snapshots, execution tuple, power/budget worksheet, schedule, and preregistration are signed and frozen.
  • Mock/oracle/executor smoke passes before paid execution.
  • Stage A completes on both frozen real repositories without unregistered changes or outcome peeking.
  • Results report the pooled confirmatory success effect, finite billed-spend endpoint, uncertainty, failures/retries, and all secondary diagnostics.
  • The result feeds #25 and triggers attribution only under #17's frozen success rule; otherwise it records a null at the frozen detection floor.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading docs/research/eval-preregistration.md and the dependency issues #35, #34, #33, and #17. Review the checked-in fake-OpenCode spike before implementing the isolated measured executor and its preparation, telemetry, retry, and analysis boundaries. Done means all acceptance checklist items pass and the frozen Stage A result is published without unregistered changes.

Written by the indexing model from the issue text.

Assessment

Domain
ai, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.