redhat-et / redhat-et/ProtoBot

Calibrate the Eliciting-Requirements Evaluation Baseline

Open
#63 5 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

component:specification-toolkit requires-manual-review
Dominant language
Go
Stars
5
Forks
6
Avg merge
23h 24m
Merged PRs (30d)
66

Description

Purpose

Complete the independent human calibration required before the Agent Eval Harness baseline for #56 is treated as trusted.

Dependency

Blocked by #56 and PR #64 landing the skill, corpus, judges, and recorded baseline.

Scope

At least two independent reviewers assess a documented sample from both development and regression partitions against property-based references, not exact wording. The sample must cover all six EARS patterns, complex-pattern splitting, false-ready rejection, both readiness contracts, consistency analysis, implementation leakage, uncertainty, host independence, and controlled-language findings.

Deliverables

  • Reviewer scores and critical-failure flags per sampled case.
  • Agreement analysis with deterministic and semantic judges.
  • Adjudication notes identifying rubric, judge, skill, or corpus corrections.
  • A committed calibration artifact with reviewer, date, model/settings, harness revision, thresholds, cost, and latency metadata.
  • Confirmed failures added to the held-out regression corpus before changing the skill or promoting a new baseline.

Acceptance criteria

  • The sample selection and reviewer assignments are recorded before scoring.
  • Any false-ready result is a critical failure regardless of aggregate score.
  • The prior baseline remains available and per-case evidence is retained.
  • The baseline is not marked trusted until calibration and adjudication complete.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading issue #56 and PR #64, which are prerequisites for the skill, corpus, judges, and recorded baseline. After they land, review the development and regression partitions and record sample selection and reviewer assignments before scoring. Done means the calibration artifact, agreement and adjudication evidence, retained prior baseline, and confirmed regression failures are committed before trust is promoted.

Written by the indexing model from the issue text.

Assessment

Domain
testing-qa
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.