callstack / callstack/agent-device

Parser fuzz validation targets: findings ledger and kill criterion (#1781 B2)

Aperta Adatta ai principianti
#1,869 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
TypeScript
Stelle
4.6k
Fork
299
Merge medio
10h 42m
PR unite (30g)
493

Descrizione

The parser fuzz lane's validation targets (`cli-validation`, `maestro-validation`, added in #1866 for #1781 B2) ship with a kill criterion that nothing currently records, so this is their live tracker — the same treatment #1823 gave the `subprocess-stub` project.

## What the lane costs

Nightly `Parser Fuzz Lane` in `.github/workflows/replays-nightly.yml`: 38,000 cases x 7 targets, measured at parity with the pre-B2 five-target/50k budget (16.9s vs 16.8s on a quiet host; the CI step was 21s before). No new job, no macOS occupancy. PR-time cost is `scripts/fuzz/validation-arbitraries.test.ts` (~0.4s, unit-core) plus the corpus replay that already existed.

## Findings ledger (append one row per nightly finding)

| date | run | target · class | real defect or phantom? | outcome |
| --- | --- | --- | --- | --- |
| _(none yet)_ | | | | |

A **phantom** is a case whose *expectation* was wrong rather than the parser — a generator drift, not a bug. One was caught pre-merge (`--scale=1.110000000000017`, float modulo drifting past a fractional `max`) and fixed before landing; `validation-arbitraries.test.ts` exists to catch that class at PR time.

## Kill criteria (either fires ⇒ delete the two validation targets, keep the classic five)

1. **Phantoms:** a phantom finding reaches a nightly twice.
2. **Yield:** no real defect found by either validation target in 6 months of nightlies — review on **2027-02-19**.

The calibration behind the lane (#1781, [B3 comment](https://github.com/callstack/agent-device/issues/1781#issuecomment-5338337202)) measured *reach*, not yield: it proved the targets can rediscover seeded defects of the shape they aim at, and explicitly did not predict how many unknown defects they will find. Criterion 2 is what tests yield, which is why this ledger exists.

## Notes for whoever reviews this

- Only two of the eight calibration rows were genuine generated new reach (#1433 via `cli-validation`; a silently-accepted Maestro field via `maestro-validation`); two more were new *detection class* but rediscovered by pinned seed cases, and four were already within the classic targets' reach. Judge yield against the first two.
- Adding a mutation class is a few lines in `scripts/fuzz/validation-arbitraries.ts`; classes whose whole input space is a handful of strings belong in the target's seed list instead.

Umbrella: #1781 (item B2/B3). Lane: #1414. Sibling tracker: #1823.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Start with .github/workflows/replays-nightly.yml and scripts/fuzz/validation-arbitraries.test.ts to understand the Parser Fuzz Lane and its validation targets. Review nightly findings and append one ledger row per finding, identifying real defects versus phantoms. Done means the tracker records each finding and the stated phantom or six-month yield criteria can be applied at review.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
github-actions, typescript
Ambito
ci-cd, documentation, testing
Tipo di issue
Documentazione
Difficoltà
2/5
Tempo stimato
1-3 ore
Stato di attività
Attiva
Chiarezza
Specificata chiaramente
Idoneità per principianti
68/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.