koala73 / koala73/worldmonitor

test(company-monitoring): build the progressive blind evaluation corpus and scorer

Open
#6,006 0 comments 0 reactions 0 assignees View on GitHub
agent-readiness area: AI/intel High Value P1
Dominant language
TypeScript
Stars
86.6k
Forks
13.1k
Avg merge
8h 4m
Merged PRs (30d)
825

Description

## Parent

#6002

## What to build

Build the versioned blind evaluation corpus and deterministic scorer used to measure discovery, attribution, materiality, direction, confidence calibration, latency, and customer usefulness through Stage 3 without pulling the post-v1 500-example expansion into this epic.

## Acceptance criteria

- [ ] Consume only the approved #6003 protocol; missing, changed, or unapproved thresholds fail before scoring and no local defaults or reinterpretation are allowed.
- [ ] Freeze 100 blind examples for tracer development and at least 200 for the Stage 3 paid-beta gate.
- [ ] Before freeze, report a non-gating feasibility forecast against all simultaneous floors: 100 published decisions for each precision/attribution gate, 75 correctly attributed material impacts overall, and 25 each for positive, negative, and mixed direction.
- [ ] Construct the 200-example candidate strata for roughly half publication-eligible examples and a sufficiently balanced eligible direction mix.
- [ ] Estimate realized rates only from a disjoint, version-locked pilot; pilot and gate sets are disjoint by occurrence, content fingerprint, company/corporate family, and source origin.
- [ ] A curator may use sealed gold labels, but classifier and policy authors see only aggregate forecasts until the release decision; curator inputs and access are versioned.
- [ ] A weak forecast returns non-blocking forecast_warning with denominator gaps and recommended untouched-example growth.
- [ ] After scoring, any denominator shortfall is blocking incomplete, not fail. Continuation retains all scored examples, appends only a precommitted untouched expansion, and rescores the cumulative corpus under the unchanged policy.
- [ ] Reject candidate-gate predictions as forecast inputs, pilot/gate overlap, policy-version mismatch, corpus mutation after lock, dropping scored examples, and fresh-corpus retries under unchanged policy.
- [ ] Score reports are deterministic and include protocol, corpus, policy/model/query versions, forecast, observed denominators, point estimates, exact bounds, calibration, confusion matrices, latency, cost, and pass/incomplete/fail.
- [ ] The Stage 4 500-example expansion is documented as a separate post-v1 release input and does not block this epic.

## Blocked by

- #6003
- #6004

## Coordination

This issue grows alongside provider and classifier implementation. Its frozen discovery protocol is a checkpoint for Exa evaluation; its applicable scorer stage must pass before classifier admission is considered complete.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the approved protocol in #6003 and the dependency in #6004. Define the versioned blind corpus and deterministic scorer around the listed acceptance criteria, including disjoint pilot and gate sets. Done means reports include the required forecasts, denominators, estimates, bounds, calibration, latency, cost, and pass/incomplete/fail status without mutating or dropping scored examples.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.