aws-samples / aws-samples/sample-autonomous-cloud-coding-agents

API: Verification evidence bundles on task completion

Abierto
#516 0 comentarios 0 reacciones 1 asignado Reclamado por @nizar-lahlali Ver en GitHub
enhancement orchestration P1
Lenguaje dominante
TypeScript
Estrellas
146
Forks
46
Merge medio
3 d 10 h
PR fusionados (30 d)
24

Descripción

> **Roadmap:** Evaluation pipeline; Tiered validation pipeline (complementary)
> **Priority:** P1

## Component

API or orchestration

## Describe the feature

Persist a structured **verification evidence bundle** on each completed (or failed) task describing which checks ran, their outcomes, and the **scope** of what was verified. Addresses the "oracle adequacy" gap from harness-engineering literature (arXiv:2605.18747 §5.1.1, §5.2.1): passing visible tests does not prove mergeability, security, or convention compliance.

Today build/lint results appear in the PR body and logs, but there is no machine-readable, API-exposed report tying verification sensors to task state.

## Use case

- **Operators** need to audit why a task succeeded or failed without reading raw CloudWatch logs.
- **Evaluation pipeline** needs structured inputs for harness-level metrics (false acceptance rate, verifier coverage).
- **Security reviewers** need explicit declaration of which gates fired (Cedar, guardrails, tests) per task.

## Proposed solution

1. **Schema** — Add `verification_report` to `TaskDetail` in `cdk/src/handlers/shared/types.ts` (mirror `cli/src/types.ts`):

```typescript
interface VerificationEvidence {
sensor: string; // e.g. "verify_build", "verify_lint", "bedrock_guardrail", "cedar_hitl"
outcome: "pass" | "fail" | "skip" | "not_applicable";
scope?: string; // human-readable limits, e.g. "unit tests only; integration tests not run"
artifact_uri?: string; // S3 link to full log
timestamp: string;
}

interface VerificationReport {
checks: VerificationEvidence[];
summary: string;
limitations?: string; // aggregate oracle adequacy note
}
```

2. **Population** — Agent post-hooks (`agent/src/post_hooks.py`) and orchestrator finalization assemble the bundle from known sensors.

3. **Surfacing** — Expose via `GET /v1/tasks/{id}`, `bgagent status --json`, and a concise section in the PR body template.

4. **Tests** — Unit tests for report assembly from fixture hook outputs.

### Acceptance criteria

- [ ] `verification_report` on `TaskDetail` with schema synced to CLI types
- [ ] At minimum: build, lint, and terminal status sensors populated for coding workflows
- [ ] PR body references the report summary
- [ ] `cdk/test/` coverage for handler assembly; agent tests for post-hook contribution
- [ ] Documented in `API_CONTRACT.md` and `EVALUATION.md`

## Other information

- **Related roadmap:** Evaluation pipeline, Tiered validation pipeline (this issue adds metadata; tiers add new checks later)
- **Paper reference:** arXiv:2605.18747 — evidence bundles as "epistemically aware" verification
- **Predecessor (soft):** Docs PEV spec issue (conceptual alignment)
- **Breaking change:** No (additive API field)

## Acknowledgements

- [ ] I may be able to implement this feature
- [ ] This might be a breaking change

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.