traverse-framework / traverse-framework/registry

Empirical evaluation of LLM-assisted capability authoring (deferred child of #355)

Open
#371 0 comments 0 reactions 0 assignees View on GitHub
enhancement quality
Dominant language
Rust
Stars
1
Forks
1
Avg merge
1h 17m
Merged PRs (30d)
217

Description

Deferred from #355 via `/brainstorm 341 355 365` (decision-log entry 85); #355 is now **closed** — its policy half shipped as approved, CI-enforced spec `023-authoring-assurance` (PRs #370/#374). This issue is the standalone owner of the empirical half. Activation trigger revised in decision-log entry 89 (`/brainstorm blocked tickets`, 2026-09-08).

## Why deferred

Nothing is blocked on the evaluation. The policy (risk-tier stance, `authoring.method` provenance marker, split reviewer qualification — all in spec `023`) gives reviewers a boundary today. The corpus + controlled run + measurement report is large work whose shape depends on there being enough real LLM-assisted publishes to measure.

## Activation condition

Start this work when **either**:

1. **≥ 5 published capabilities carry `authoring.method: "llm-assisted"`** — a one-line count over `capabilities/**/contract.json` on `main`. Progress: **~3 / 5** as of 2026-09-08 (`audio.calibration-plan-create@1.0.0`, `inference.evidence-normalize@1.0.0`, `inference.evidence-normalize@1.0.1`). `#369` was the first.
2. **or** there is an owner decision to move `023`'s low-risk deterministic tier from `experimental` to `supported` (this evaluation's data is the precondition for that move).

Activation means "begin building the corpus and methodology" — not "there is already enough data for a credible report". The report's own credibility bar (sample size, per-bucket `n`) is set by the methodology when it is written, the same split `#365` uses.

## Scope when activated (absorbs every unfinished #355 DoD item)

- A versioned evaluation corpus covering deterministic/pure, validation, effectful, and explicitly-excluded high-risk capability classes.
- Every generated candidate linked to its contract, source revision, artifact digest, test evidence, and human review decision (spec `023` FR-003 provides most of this chain).
- Report measures: first-pass validation rate, security/policy rejection rate, reviewer rework, defect escape, time-to-accepted artifact. Methodology and limitations published.
- Contract-derived tests + at least one property-based or metamorphic check where the contract admits them.
- The report's data sets the numeric bar for the `experimental` → `supported` promotion of the low-risk tier (a later standalone owner decision, per `023` FR-004 / Governing Relationship).

(#355 DoD items "policy by risk tier", "non-Rust vs qualified review guidance", and "no bypass of existing gates" are already delivered in spec `023` FR-004 / FR-005 / FR-006 and are not repeated here.)

## Definition of done

- [ ] Evaluation corpus published (private/controlled per #355 D2; publishable aggregate methodology + results).
- [ ] Report published under `docs/` with the five metrics above, sample size, and limitations.
- [ ] A `docs/decision-log.md` entry recording whether the low-risk tier is promoted to `supported`, with the data behind it.
- [ ] `023` amended (or a successor spec) if the evidence changes the tier stance.

## Related

- Was: child of #355 (now closed, superseded by spec `023` + this issue)
- Policy spec: `023-authoring-assurance`
- Trigger rationale: decision-log entry 89. Prior art for "defer the build until real demand": entries 34c/d, 83 Q3.

Contributor guide

Open the contributing guide

Research direction

Wait for the activation condition, then count `authoring.method: "llm-assisted"` entries under `capabilities/**/contract.json` and read spec `023-authoring-assurance` plus decision-log entry 89. Build the versioned corpus and methodology around the five listed metrics, linking candidates to contracts, revisions, digests, tests, and review decisions. Done means publishable aggregate results under `docs/`, a decision-log entry, and any required amendment to `023`.

Written by the indexing model from the issue text.

Assessment

Domain
ai, data, documentation, testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.