aws-samples / aws-samples/sample-autonomous-cloud-coding-agents

chore(agent): expose or remove the _debug_cw_failures counter — incremented, never read

Open
#810 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
143
Forks
46
Avg merge
3d 9h
Merged PRs (30d)
20

Description

## Problem

`agent/src/server.py` maintains a module-global `_debug_cw_failures` counter, incremented under a lock on every failed CloudWatch write (`:197`, `:226`). Nothing ever reads it.

Three docstrings describe it as an operator-facing signal — `:145` and `:176` call it *"a single alarm surface"* — but there is no metric emission, no CloudWatch alarm, and no `/ping` or `/validate` exposure. The value dies with the process.

## Why it matters

The counter exists to answer "is the debug/diagnostic path blind?" — i.e. is the agent failing to write the very lines an operator would use to diagnose a task. That is a genuinely useful signal on a substrate where the guest is otherwise unobservable, and it is currently unobtainable. Worse, the docstrings promise it, so a reader debugging a silent task may go looking for an alarm that does not exist.

Surfaced during the ADR-021 P2 review (PR #733, non-blocking item 6). Pre-existing, not introduced by that PR. Note that #733 removed the one place that *depended* on the counter rhetorically: `_build_hook_log`'s reason #1 previously argued a build-role write would "poison the signal", and now argues from IAM namespace scoping instead. So there is no longer anything blocking a decision either way.

## Options

1. **Expose it** (smallest useful change): include it in the `/ping` and/or `/validate` response body. `/validate` already has a `warnings` array and, as of #733, logs named warnings to the build log group — the same shape would work here.
2. **Emit it as a metric** (most useful): an EMF line or a `PutMetricData` call at task finalize, which gives operators something alarmable. Note the ordering constraint: the emitter must not itself be a CloudWatch write on the path it is measuring, and it must respect the pre-`platform_config` AWS-silence rule (see `_aws_silent_log`).
3. **Remove it** and correct the three docstrings. Legitimate if nobody wants the signal — the failures are already logged individually to stdout.

## Acceptance criteria

- `_debug_cw_failures` is either read by something an operator can observe, or removed.
- The docstrings at `server.py:145` and `:176` describe what actually exists.
- If exposed: a test asserts a non-zero counter is visible after a forced write failure, and that the exposure path does not itself perform a CloudWatch write.
- If removed: no docstring anywhere references an alarm surface.

Contributor guide

Open the contributing guide

Research direction

Start in agent/src/server.py around the _debug_cw_failures increments at lines 197 and 226, then inspect the /ping and /validate responses and the _aws_silent_log ordering rule. Decide whether the counter should be exposed, emitted, or removed, and verify the relevant failure path and docstrings. Done means the counter's behavior matches the chosen option and no documentation promises an unavailable alarm surface.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python
Domain
backend, cloud, observability
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.