redhat-et / redhat-et/ProtoBot
Calibrate the Eliciting-Requirements Evaluation Baseline
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 5
- Forks
- 6
- Avg merge
- 23h 24m
- Merged PRs (30d)
- 66
Description
Purpose
Complete the independent human calibration required before the Agent Eval Harness baseline for #56 is treated as trusted.
Dependency
Blocked by #56 and PR #64 landing the skill, corpus, judges, and recorded baseline.
Scope
At least two independent reviewers assess a documented sample from both development and regression partitions against property-based references, not exact wording. The sample must cover all six EARS patterns, complex-pattern splitting, false-ready rejection, both readiness contracts, consistency analysis, implementation leakage, uncertainty, host independence, and controlled-language findings.
Deliverables
- Reviewer scores and critical-failure flags per sampled case.
- Agreement analysis with deterministic and semantic judges.
- Adjudication notes identifying rubric, judge, skill, or corpus corrections.
- A committed calibration artifact with reviewer, date, model/settings, harness revision, thresholds, cost, and latency metadata.
- Confirmed failures added to the held-out regression corpus before changing the skill or promoting a new baseline.
Acceptance criteria
- The sample selection and reviewer assignments are recorded before scoring.
- Any false-ready result is a critical failure regardless of aggregate score.
- The prior baseline remains available and per-case evidence is retained.
- The baseline is not marked trusted until calibration and adjudication complete.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading issue #56 and PR #64, which are prerequisites for the skill, corpus, judges, and recorded baseline. After they land, review the development and regression partitions and record sample selection and reviewer assignments before scoring. Done means the calibration artifact, agreement and adjudication evidence, retained prior baseline, and confirmed regression failures are committed before trust is promoted.
Written by the indexing model from the issue text.
Assessment
- Domain
- testing-qa
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 35/100