aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
chore(agent): expose or remove the _debug_cw_failures counter — incremented, never read
- Lingua principale
- TypeScript
- Stelle
- 143
- Fork
- 46
- Merge medio
- 3g 9h
- PR unite (30g)
- 20
Descrizione
## Problem
`agent/src/server.py` maintains a module-global `_debug_cw_failures` counter, incremented under a lock on every failed CloudWatch write (`:197`, `:226`). Nothing ever reads it.
Three docstrings describe it as an operator-facing signal — `:145` and `:176` call it *"a single alarm surface"* — but there is no metric emission, no CloudWatch alarm, and no `/ping` or `/validate` exposure. The value dies with the process.
## Why it matters
The counter exists to answer "is the debug/diagnostic path blind?" — i.e. is the agent failing to write the very lines an operator would use to diagnose a task. That is a genuinely useful signal on a substrate where the guest is otherwise unobservable, and it is currently unobtainable. Worse, the docstrings promise it, so a reader debugging a silent task may go looking for an alarm that does not exist.
Surfaced during the ADR-021 P2 review (PR #733, non-blocking item 6). Pre-existing, not introduced by that PR. Note that #733 removed the one place that *depended* on the counter rhetorically: `_build_hook_log`'s reason #1 previously argued a build-role write would "poison the signal", and now argues from IAM namespace scoping instead. So there is no longer anything blocking a decision either way.
## Options
1. **Expose it** (smallest useful change): include it in the `/ping` and/or `/validate` response body. `/validate` already has a `warnings` array and, as of #733, logs named warnings to the build log group — the same shape would work here.
2. **Emit it as a metric** (most useful): an EMF line or a `PutMetricData` call at task finalize, which gives operators something alarmable. Note the ordering constraint: the emitter must not itself be a CloudWatch write on the path it is measuring, and it must respect the pre-`platform_config` AWS-silence rule (see `_aws_silent_log`).
3. **Remove it** and correct the three docstrings. Legitimate if nobody wants the signal — the failures are already logged individually to stdout.
## Acceptance criteria
- `_debug_cw_failures` is either read by something an operator can observe, or removed.
- The docstrings at `server.py:145` and `:176` describe what actually exists.
- If exposed: a test asserts a non-zero counter is visible after a forced write failure, and that the exposure path does not itself perform a CloudWatch write.
- If removed: no docstring anywhere references an alarm surface.
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia in agent/src/server.py intorno agli incrementi di _debug_cw_failures alle righe 197 e 226, quindi esamina le risposte /ping e /validate e la regola di ordinamento di _aws_silent_log. Decidi se il contatore debba essere esposto, emesso o rimosso, e verifica il percorso di errore pertinente e i docstrings. L’attività è completata quando il comportamento del contatore corrisponde all’opzione scelta e nessuna documentazione promette una superficie di allarme non disponibile.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- aws, python
- Ambito
- backend, cloud, observability
- Tipo di issue
- Refactoring
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 45/100