P13 — Agent-outcome evaluation: measure matlatl on documentation-dependent coding tasks
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Why
Every agent-facing claim in matlatl remains unvalidated. Generated repository context can improve navigation, add cost without improving completion, or harm task success. We need controlled evidence before changing defaults or funding more analytics.
This issue now governs Experiment A:
On a frozen benchmark of real coding-change tasks whose correct implementation depends on repository documentation, what is matlatl's marginal effect on deterministic hidden-verifier success and billed cost?
This does not answer whether healthy documentation itself causally helps coding agents. That is a separate future Experiment B requiring controlled healthy/degraded documentation interventions while holding code and task fixed. Adopted versus unadopted repositories are descriptive, not causal evidence.
Current state
Completed:
- Go-only offline scaffold (#32).
- 97-case deterministic Level 1 correctness oracle (#34).
- A real offline smevals + pinned OpenCode boundary spike with fake OpenCode, immutable records, private post-run scoring, and remote streamable-HTTP matlatl MCP.
Not completed:
- isolated measured executor and telemetry;
- source-bearing Nimbus coding fixtures (#35);
- model/provider qualification;
- frozen real-corpus task families and signed budget/power design;
- paid Stage A execution (#36).
Binding methodology:
docs/research/eval-overview.mddocs/research/eval-harness-design.mddocs/research/eval-preregistration.md
This body supersedes the earlier Inspect AI/Claude Code, four-arm, documentation-task design.
Experiment A design
Tasks
Every measured task requires a code edit and deterministic private verification of a documented behavioral, security, configuration, or architecture constraint. The mutually exclusive primary strata are:
- grep-friendly coding control;
- cross-document synthesis;
- navigation-heavy coding;
- single-document constraint.
Tasks are grouped into independent task families. The estimand is the average paired effect on the frozen two-repository benchmark; external validity is limited accordingly. Navigation-only QA, comprehension QA, and documentation repair are calibration or focused studies, not headline evidence.
Arms
The first paid Stage A has two arms only:
baseline: normalized frozen snapshot with pre-existing matlatl artifacts/MCP removed;+all(machine IDall): identical snapshot plus freshly generated rootllms.txt, roottrails.json, remote streamable-HTTP matlatl MCP, and one frozen availability notice.
Pointer/trails/llms/MCP attribution is a separately preregistered study. It is triggered automatically only by the frozen pooled success statistic, or funded and signed separately after a null.
Selection, telemetry, and analysis
- Select the pinned OpenCode model/provider using baseline-only disposable Nimbus coding tasks and frozen competence, protocol reliability, telemetry, and budget criteria. Never select on treatment response.
- Disable caching. Qualification must prove explicit zero cache-read/write counters.
- Record uncached input/output tokens, externally reconciled billed cost, wall time, turns, total tool calls,
llms.txt/trails.jsonaccesses, and calls to each matlatl MCP tool. - The first model request is the exposure boundary. Only pre-exposure infrastructure failures may receive bounded fresh retries. Post-exposure failures remain assigned-arm outcomes in the primary intention-to-treat success analysis.
- Primary confirmatory statistic: pooled task-family success contrast,
+all - baseline. - Primary finite cost endpoint: task-family mean billed spend per scheduled attempt, including unsuccessful attempts and retry spend; missing post-exposure billed cost receives the frozen conservative cap.
- Cost, corpus, stratum, and mechanism telemetry are prespecified secondary/descriptive results and cannot trigger attribution.
Power and schedule
Task families N and repetitions r are selected before freeze by a signed budget/power calculation over feasible 2 × N × r designs. The calculation models the exact success/cost estimands, favors more independent task families when power is near-equal, starts comparison at r = 2, reserves qualification/retry budget, and raises the disclosed detection floor if the target is unaffordable.
Delivery order
- Isolated executor and normalized arm preparation with fake probes.
- Nimbus source/code fixtures and mechanical coverage (#35).
- Zero-cache/access telemetry, retry/exposure accounting, and baseline-only qualification substrate.
- Real-corpus eligible frames, reviewed task families, immutable snapshots, signed budget/power worksheet, and signed preregistration.
- Paid two-arm Stage A (#36), without outcome peeking.
- Conditional attribution or separately funded Experiment B.
#33 remains a separate signal-quality workstream and does not enter Experiment A's endpoint or model selection.
Done
- Two frozen real repositories and independently reviewed coding-task families are prepared.
- The complete execution tuple, zero-cache proof, eligible-frame selection, telemetry mapping, failure rules, schedule, and signed budget/power design are frozen before measurement.
- Stage A baseline and
+allruns complete through the isolated smevals/OpenCode executor. - Results publish the pooled confirmatory success contrast plus billed spend, failures, retries, telemetry, corpus/stratum summaries, uncertainty, and the finite-benchmark limitation.
- The result feeds #25 and either triggers a separately preregistered attribution study or records a null at the frozen detection floor.
A flat result means “no effect detected at this frozen floor,” not equivalence.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading docs/research/eval-overview.md, docs/research/eval-harness-design.md, and docs/research/eval-preregistration.md, then inspect the completed offline scaffold and the smevals/OpenCode boundary spike. Trace the missing isolated executor, telemetry, fixtures, qualification, and signed budget/power work in the stated delivery order. Done means the frozen two-arm benchmark runs and publishes the listed success, cost, failure, telemetry, and limitation results.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go
- Domain
- ai-infra-agents, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100