aws-samples / aws-samples/sample-autonomous-cloud-coding-agents

chore(agent): expose or remove the _debug_cw_failures counter — incremented, never read

Abierto
#810 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
TypeScript
Estrellas
143
Forks
46
Merge medio
3 d 10 h
PR fusionados (30 d)
24

Descripción

## Problem

`agent/src/server.py` maintains a module-global `_debug_cw_failures` counter, incremented under a lock on every failed CloudWatch write (`:197`, `:226`). Nothing ever reads it.

Three docstrings describe it as an operator-facing signal — `:145` and `:176` call it *"a single alarm surface"* — but there is no metric emission, no CloudWatch alarm, and no `/ping` or `/validate` exposure. The value dies with the process.

## Why it matters

The counter exists to answer "is the debug/diagnostic path blind?" — i.e. is the agent failing to write the very lines an operator would use to diagnose a task. That is a genuinely useful signal on a substrate where the guest is otherwise unobservable, and it is currently unobtainable. Worse, the docstrings promise it, so a reader debugging a silent task may go looking for an alarm that does not exist.

Surfaced during the ADR-021 P2 review (PR #733, non-blocking item 6). Pre-existing, not introduced by that PR. Note that #733 removed the one place that *depended* on the counter rhetorically: `_build_hook_log`'s reason #1 previously argued a build-role write would "poison the signal", and now argues from IAM namespace scoping instead. So there is no longer anything blocking a decision either way.

## Options

1. **Expose it** (smallest useful change): include it in the `/ping` and/or `/validate` response body. `/validate` already has a `warnings` array and, as of #733, logs named warnings to the build log group — the same shape would work here.
2. **Emit it as a metric** (most useful): an EMF line or a `PutMetricData` call at task finalize, which gives operators something alarmable. Note the ordering constraint: the emitter must not itself be a CloudWatch write on the path it is measuring, and it must respect the pre-`platform_config` AWS-silence rule (see `_aws_silent_log`).
3. **Remove it** and correct the three docstrings. Legitimate if nobody wants the signal — the failures are already logged individually to stdout.

## Acceptance criteria

- `_debug_cw_failures` is either read by something an operator can observe, or removed.
- The docstrings at `server.py:145` and `:176` describe what actually exists.
- If exposed: a test asserts a non-zero counter is visible after a forced write failure, and that the exposure path does not itself perform a CloudWatch write.
- If removed: no docstring anywhere references an alarm surface.

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Empieza en agent/src/server.py, alrededor de los incrementos de _debug_cw_failures en las líneas 197 y 226; después inspecciona las respuestas de /ping y /validate y la regla de ordenación de _aws_silent_log. Decide si el contador debe exponerse, emitirse o eliminarse, y verifica la ruta de fallo correspondiente y los docstrings. Se considera terminado cuando el comportamiento del contador coincide con la opción elegida y ninguna documentación promete una superficie de alarmas no disponible.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws, python
Área
backend, cloud, observability
Tipo de issue
Refactorización
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Activo
Claridad
Bastante claro
Aptitud para principiantes
45/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.