traverse-framework / traverse-framework/registry
Empirical evaluation of LLM-assisted capability authoring (deferred child of #355)
- Dominant language
- Rust
- Stars
- 1
- Forks
- 1
- Avg merge
- 1h 17m
- Merged PRs (30d)
- 217
Description
Deferred from #355 via `/brainstorm 341 355 365` (decision-log entry 85); #355 is now **closed** — its policy half shipped as approved, CI-enforced spec `023-authoring-assurance` (PRs #370/#374). This issue is the standalone owner of the empirical half. Activation trigger revised in decision-log entry 89 (`/brainstorm blocked tickets`, 2026-09-08).
## Why deferred
Nothing is blocked on the evaluation. The policy (risk-tier stance, `authoring.method` provenance marker, split reviewer qualification — all in spec `023`) gives reviewers a boundary today. The corpus + controlled run + measurement report is large work whose shape depends on there being enough real LLM-assisted publishes to measure.
## Activation condition
Start this work when **either**:
1. **≥ 5 published capabilities carry `authoring.method: "llm-assisted"`** — a one-line count over `capabilities/**/contract.json` on `main`. Progress: **~3 / 5** as of 2026-09-08 (`audio.calibration-plan-create@1.0.0`, `inference.evidence-normalize@1.0.0`, `inference.evidence-normalize@1.0.1`). `#369` was the first.
2. **or** there is an owner decision to move `023`'s low-risk deterministic tier from `experimental` to `supported` (this evaluation's data is the precondition for that move).
Activation means "begin building the corpus and methodology" — not "there is already enough data for a credible report". The report's own credibility bar (sample size, per-bucket `n`) is set by the methodology when it is written, the same split `#365` uses.
## Scope when activated (absorbs every unfinished #355 DoD item)
- A versioned evaluation corpus covering deterministic/pure, validation, effectful, and explicitly-excluded high-risk capability classes.
- Every generated candidate linked to its contract, source revision, artifact digest, test evidence, and human review decision (spec `023` FR-003 provides most of this chain).
- Report measures: first-pass validation rate, security/policy rejection rate, reviewer rework, defect escape, time-to-accepted artifact. Methodology and limitations published.
- Contract-derived tests + at least one property-based or metamorphic check where the contract admits them.
- The report's data sets the numeric bar for the `experimental` → `supported` promotion of the low-risk tier (a later standalone owner decision, per `023` FR-004 / Governing Relationship).
(#355 DoD items "policy by risk tier", "non-Rust vs qualified review guidance", and "no bypass of existing gates" are already delivered in spec `023` FR-004 / FR-005 / FR-006 and are not repeated here.)
## Definition of done
- [ ] Evaluation corpus published (private/controlled per #355 D2; publishable aggregate methodology + results).
- [ ] Report published under `docs/` with the five metrics above, sample size, and limitations.
- [ ] A `docs/decision-log.md` entry recording whether the low-risk tier is promoted to `supported`, with the data behind it.
- [ ] `023` amended (or a successor spec) if the evidence changes the tier stance.
## Related
- Was: child of #355 (now closed, superseded by spec `023` + this issue)
- Policy spec: `023-authoring-assurance`
- Trigger rationale: decision-log entry 89. Prior art for "defer the build until real demand": entries 34c/d, 83 Q3.
Contributor guide
Research direction
Wait for the activation condition, then count `authoring.method: "llm-assisted"` entries under `capabilities/**/contract.json` and read spec `023-authoring-assurance` plus decision-log entry 89. Build the versioned corpus and methodology around the five listed metrics, linking candidates to contracts, revisions, digests, tests, and review decisions. Done means publishable aggregate results under `docs/`, a decision-log entry, and any required amendment to `023`.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, data, documentation, testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 32/100