aws-samples / aws-samples/sample-autonomous-cloud-coding-agents

CDK: Harness-level SLO metrics on operator dashboard

Offen
#511 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
enhancement infra-cdk P1
Vorherrschende Sprache
TypeScript
Sterne
143
Forks
46
Ø Merge
3 T. 9 Std.
Gemergte PRs (30 T.)
20

Beschreibung

> **Roadmap:** Evaluation pipeline; Validation and risk analytics
> **Priority:** P1

## Component

CDK / infrastructure

## Describe the feature

Emit **harness-level** CloudWatch metrics and dashboard panels that evaluate the control plane/runtime separately from raw task pass/fail. Final task success conflates model quality, harness quality, and environment luck (arXiv:2605.18747 §5.2.1).

## Use case

- **Operators/SREs** track whether the platform is getting more reliable independent of model upgrades.
- **Field engagements** report dark-factory scorecard progress with objective harness SLOs.
- **A/B prompt experiments** isolate prompt changes from harness regressions.

## Proposed solution

Define and emit metrics (names illustrative; finalize in implementation):

| Metric | Definition |
|--------|------------|
| `HarnessToolCallsPerTask` | Agent tool invocations per completed task |
| `HarnessTokensPerTask` | Bedrock tokens (from existing cost telemetry) |
| `HarnessWallClockSeconds` | Submit → terminal |
| `HarnessVerificationCoverage` | % tasks with full evidence bundle / all required sensors run |
| `HarnessRecoveryRate` | % tasks that failed verify then succeeded on retry (when fix-up loop exists) |
| `HarnessReplayCompleteness` | % tasks with TaskEvents + trace URI + prompt_version |

Add widgets to the existing operator dashboard construct; document dimensions (`repo`, `workflow_ref`, `prompt_version`) without PII.

### Acceptance criteria

- [ ] Metrics emitted from orchestrator and/or agent telemetry path
- [ ] Dashboard panels added to operator dashboard (CDK)
- [ ] Documented in `docs/design/OBSERVABILITY.md` with metric definitions
- [ ] Referenced from dark-factory scorecard / `ABCA_V2.md` success metrics where appropriate
- [ ] No new PII dimensions in metric labels

## Other information

- **Depends on (soft):** Verification evidence bundles for `VerificationCoverage`
- **Paper:** arXiv:2605.18747 §5.2.1 harness-level evaluation dimensions
- **Existing:** Operator dashboard, OTEL spans, TaskEvents (shipped)

## Acknowledgements

- [ ] I may be able to implement this feature
- [ ] This might be a breaking change

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit dem bestehenden Operator-Dashboard-Konstrukt und dem Telemetriepfad des Orchestrators oder Agents und überprüfe dann die vorhandenen OTEL-Spans, TaskEvents und Kostentelemetrie. Definiere die Harness-Metriken und nicht personenbezogenen Dimensionen, füge Dashboard-Panels hinzu, dokumentiere sie in docs/design/OBSERVABILITY.md und aktualisiere gegebenenfalls die Erfolgsmetriken des dark-factory-Scorecards oder von ABCA_V2.md.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
aws, typescript
Bereich
cloud, infrastructure, observability-sre
Issue-Typ
Feature
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
48/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.