koala73 / koala73/worldmonitor
test(company-monitoring): build the progressive blind evaluation corpus and scorer
- Dominant language
- TypeScript
- Stars
- 86.6k
- Forks
- 13.1k
- Avg merge
- 8h 4m
- Merged PRs (30d)
- 825
Description
## Parent
#6002
## What to build
Build the versioned blind evaluation corpus and deterministic scorer used to measure discovery, attribution, materiality, direction, confidence calibration, latency, and customer usefulness through Stage 3 without pulling the post-v1 500-example expansion into this epic.
## Acceptance criteria
- [ ] Consume only the approved #6003 protocol; missing, changed, or unapproved thresholds fail before scoring and no local defaults or reinterpretation are allowed.
- [ ] Freeze 100 blind examples for tracer development and at least 200 for the Stage 3 paid-beta gate.
- [ ] Before freeze, report a non-gating feasibility forecast against all simultaneous floors: 100 published decisions for each precision/attribution gate, 75 correctly attributed material impacts overall, and 25 each for positive, negative, and mixed direction.
- [ ] Construct the 200-example candidate strata for roughly half publication-eligible examples and a sufficiently balanced eligible direction mix.
- [ ] Estimate realized rates only from a disjoint, version-locked pilot; pilot and gate sets are disjoint by occurrence, content fingerprint, company/corporate family, and source origin.
- [ ] A curator may use sealed gold labels, but classifier and policy authors see only aggregate forecasts until the release decision; curator inputs and access are versioned.
- [ ] A weak forecast returns non-blocking forecast_warning with denominator gaps and recommended untouched-example growth.
- [ ] After scoring, any denominator shortfall is blocking incomplete, not fail. Continuation retains all scored examples, appends only a precommitted untouched expansion, and rescores the cumulative corpus under the unchanged policy.
- [ ] Reject candidate-gate predictions as forecast inputs, pilot/gate overlap, policy-version mismatch, corpus mutation after lock, dropping scored examples, and fresh-corpus retries under unchanged policy.
- [ ] Score reports are deterministic and include protocol, corpus, policy/model/query versions, forecast, observed denominators, point estimates, exact bounds, calibration, confusion matrices, latency, cost, and pass/incomplete/fail.
- [ ] The Stage 4 500-example expansion is documented as a separate post-v1 release input and does not block this epic.
## Blocked by
- #6003
- #6004
## Coordination
This issue grows alongside provider and classifier implementation. Its frozen discovery protocol is a checkpoint for Exa evaluation; its applicable scorer stage must pass before classifier admission is considered complete.
Contributor guide
Research direction
Start by reviewing the approved protocol in #6003 and the dependency in #6004. Define the versioned blind corpus and deterministic scorer around the listed acceptance criteria, including disjoint pilot and gate sets. Done means reports include the required forecasts, denominators, estimates, bounds, calibration, latency, cost, and pass/incomplete/fail status without mutating or dropping scored examples.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100