aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
chore(agent): expose or remove the _debug_cw_failures counter — incremented, never read
- Lenguaje dominante
- TypeScript
- Estrellas
- 143
- Forks
- 46
- Merge medio
- 3 d 10 h
- PR fusionados (30 d)
- 24
Descripción
## Problem
`agent/src/server.py` maintains a module-global `_debug_cw_failures` counter, incremented under a lock on every failed CloudWatch write (`:197`, `:226`). Nothing ever reads it.
Three docstrings describe it as an operator-facing signal — `:145` and `:176` call it *"a single alarm surface"* — but there is no metric emission, no CloudWatch alarm, and no `/ping` or `/validate` exposure. The value dies with the process.
## Why it matters
The counter exists to answer "is the debug/diagnostic path blind?" — i.e. is the agent failing to write the very lines an operator would use to diagnose a task. That is a genuinely useful signal on a substrate where the guest is otherwise unobservable, and it is currently unobtainable. Worse, the docstrings promise it, so a reader debugging a silent task may go looking for an alarm that does not exist.
Surfaced during the ADR-021 P2 review (PR #733, non-blocking item 6). Pre-existing, not introduced by that PR. Note that #733 removed the one place that *depended* on the counter rhetorically: `_build_hook_log`'s reason #1 previously argued a build-role write would "poison the signal", and now argues from IAM namespace scoping instead. So there is no longer anything blocking a decision either way.
## Options
1. **Expose it** (smallest useful change): include it in the `/ping` and/or `/validate` response body. `/validate` already has a `warnings` array and, as of #733, logs named warnings to the build log group — the same shape would work here.
2. **Emit it as a metric** (most useful): an EMF line or a `PutMetricData` call at task finalize, which gives operators something alarmable. Note the ordering constraint: the emitter must not itself be a CloudWatch write on the path it is measuring, and it must respect the pre-`platform_config` AWS-silence rule (see `_aws_silent_log`).
3. **Remove it** and correct the three docstrings. Legitimate if nobody wants the signal — the failures are already logged individually to stdout.
## Acceptance criteria
- `_debug_cw_failures` is either read by something an operator can observe, or removed.
- The docstrings at `server.py:145` and `:176` describe what actually exists.
- If exposed: a test asserts a non-zero counter is visible after a forced write failure, and that the exposure path does not itself perform a CloudWatch write.
- If removed: no docstring anywhere references an alarm surface.
Guía de contribución
Línea de trabajo
Empieza en agent/src/server.py, alrededor de los incrementos de _debug_cw_failures en las líneas 197 y 226; después inspecciona las respuestas de /ping y /validate y la regla de ordenación de _aws_silent_log. Decide si el contador debe exponerse, emitirse o eliminarse, y verifica la ruta de fallo correspondiente y los docstrings. Se considera terminado cuando el comportamiento del contador coincide con la opción elegida y ninguna documentación promete una superficie de alarmas no disponible.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- aws, python
- Área
- backend, cloud, observability
- Tipo de issue
- Refactorización
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Activo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 45/100