callstack / callstack/agent-device

Parser fuzz validation targets: findings ledger and kill criterion (#1781 B2)

Offen Anfängerfreundlich
#1,869 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
TypeScript
Sterne
4.7k
Forks
303
Ø Merge
10 Std. 42 Min.
Gemergte PRs (30 T.)
493

Beschreibung

The parser fuzz lane's validation targets (cli-validation, maestro-validation, added in #1866 for #1781 B2) ship with a kill criterion that nothing currently records, so this is their live tracker — the same treatment #1823 gave the subprocess-stub project.

What the lane costs

Nightly Parser Fuzz Lane in .github/workflows/replays-nightly.yml: 38,000 cases x 7 targets, measured at parity with the pre-B2 five-target/50k budget (16.9s vs 16.8s on a quiet host; the CI step was 21s before). No new job, no macOS occupancy. PR-time cost is scripts/fuzz/validation-arbitraries.test.ts (~0.4s, unit-core) plus the corpus replay that already existed.

Findings ledger (append one row per nightly finding)

date run target · class real defect or phantom? outcome
(none yet)

A phantom is a case whose expectation was wrong rather than the parser — a generator drift, not a bug. One was caught pre-merge (--scale=1.110000000000017, float modulo drifting past a fractional max) and fixed before landing; validation-arbitraries.test.ts exists to catch that class at PR time.

Kill criteria (either fires ⇒ delete the two validation targets, keep the classic five)

  1. Phantoms: a phantom finding reaches a nightly twice.
  2. Yield: no real defect found by either validation target in 6 months of nightlies — review on 2027-02-19.

The calibration behind the lane (#1781, B3 comment) measured reach, not yield: it proved the targets can rediscover seeded defects of the shape they aim at, and explicitly did not predict how many unknown defects they will find. Criterion 2 is what tests yield, which is why this ledger exists.

Notes for whoever reviews this

  • Only two of the eight calibration rows were genuine generated new reach (#1433 via cli-validation; a silently-accepted Maestro field via maestro-validation); two more were new detection class but rediscovered by pinned seed cases, and four were already within the classic targets' reach. Judge yield against the first two.
  • Adding a mutation class is a few lines in scripts/fuzz/validation-arbitraries.ts; classes whose whole input space is a handful of strings belong in the target's seed list instead.

Umbrella: #1781 (item B2/B3). Lane: #1414. Sibling tracker: #1823.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne mit .github/workflows/replays-nightly.yml und scripts/fuzz/validation-arbitraries.test.ts, um die Parser Fuzz Lane und ihre Validierungsziele zu verstehen. Prüfe die nächtlichen Findings und füge pro Finding eine Zeile im Ledger hinzu, in der echte Fehler von Phantomen unterschieden werden. Fertig bedeutet, dass der Tracker jedes Finding erfasst und die angegebenen Phantom- oder Sechsmonats-Yield-Kriterien bei der Prüfung angewendet werden können.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
github-actions, typescript
Bereich
ci-cd, documentation, testing
Issue-Typ
Dokumentation
Schwierigkeit
2/5
Geschätzter Aufwand
1-3 Stunden
Aktivitätsstatus
Aktiv
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
68/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.