aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
CDK: Harness-level SLO metrics on operator dashboard
- 主要言語
- TypeScript
- スター
- 143
- フォーク
- 46
- 平均マージ
- 3日 9時間
- マージ済み PR(30日)
- 20
説明
> **Roadmap:** Evaluation pipeline; Validation and risk analytics
> **Priority:** P1
## Component
CDK / infrastructure
## Describe the feature
Emit **harness-level** CloudWatch metrics and dashboard panels that evaluate the control plane/runtime separately from raw task pass/fail. Final task success conflates model quality, harness quality, and environment luck (arXiv:2605.18747 §5.2.1).
## Use case
- **Operators/SREs** track whether the platform is getting more reliable independent of model upgrades.
- **Field engagements** report dark-factory scorecard progress with objective harness SLOs.
- **A/B prompt experiments** isolate prompt changes from harness regressions.
## Proposed solution
Define and emit metrics (names illustrative; finalize in implementation):
| Metric | Definition |
|--------|------------|
| `HarnessToolCallsPerTask` | Agent tool invocations per completed task |
| `HarnessTokensPerTask` | Bedrock tokens (from existing cost telemetry) |
| `HarnessWallClockSeconds` | Submit → terminal |
| `HarnessVerificationCoverage` | % tasks with full evidence bundle / all required sensors run |
| `HarnessRecoveryRate` | % tasks that failed verify then succeeded on retry (when fix-up loop exists) |
| `HarnessReplayCompleteness` | % tasks with TaskEvents + trace URI + prompt_version |
Add widgets to the existing operator dashboard construct; document dimensions (`repo`, `workflow_ref`, `prompt_version`) without PII.
### Acceptance criteria
- [ ] Metrics emitted from orchestrator and/or agent telemetry path
- [ ] Dashboard panels added to operator dashboard (CDK)
- [ ] Documented in `docs/design/OBSERVABILITY.md` with metric definitions
- [ ] Referenced from dark-factory scorecard / `ABCA_V2.md` success metrics where appropriate
- [ ] No new PII dimensions in metric labels
## Other information
- **Depends on (soft):** Verification evidence bundles for `VerificationCoverage`
- **Paper:** arXiv:2605.18747 §5.2.1 harness-level evaluation dimensions
- **Existing:** Operator dashboard, OTEL spans, TaskEvents (shipped)
## Acknowledgements
- [ ] I may be able to implement this feature
- [ ] This might be a breaking change
コントリビューションガイド
調査の方向性
既存のオペレーターダッシュボードの構成と、オーケストレーターまたはエージェントのテレメトリーパスから始め、既存の OTEL span、TaskEvents、コストテレメトリーを確認します。ハーネスのメトリクスと非 PII のディメンションを定義し、ダッシュボードパネルを追加して、docs/design/OBSERVABILITY.md にドキュメント化し、必要に応じて dark-factory スコアカードまたは ABCA_V2.md の成功メトリクスを更新します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- aws, typescript
- 領域
- cloud, infrastructure, observability-sre
- issue の種類
- 機能追加
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 48/100