eclipse-score / eclipse-score/score

Build Internal Reliability Benchmark Suite

Open
#2,834 0 comments 0 reactions 0 assignees View on GitHub
automation safety testing
Dominant language
Starlark
Stars
109
Forks
105
Avg merge
22h 15m
Merged PRs (30d)
22

Description

## Goal
Create an internal benchmark that measures real reliability before org-wide hard enforcement.

## Why
Hard-fail governance should only be enabled after measurable, repeatable reliability in real repository tasks.

## Tasks
- Curate representative task set from real issues.
- Define scoring dimensions: success, regressions, policy violations, time-to-green.
- Split KPIs into:
- Lane A mandatory KPI set (merge-governing)
- Lane B optional accelerator delta (optimization only)
- Run benchmark on scheduled cadence.
- Publish dashboards and trend reports.

## Initial KPI Targets (proposed, calibrate after pilot)
- Lane A success rate >= 80%
- Policy-violating outcome <= 5%
- Rerun divergence <= 10%

## Done When
- Benchmark suite runs on a fixed cadence.
- KPI thresholds are approved and tied to governance decisions.
- Hard-fail rollout criteria explicitly depend on Lane A KPI thresholds.

Parent: https://github.com/eclipse-score/score/issues/2827

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.