jantimon / jantimon/web-performance-debugger
Background subagents stall when ending their turn on a pending background task
- Dominant language
- TypeScript
- Stars
- 31
- Forks
- 1
- Avg merge
- 1h 18m
- Merged PRs (30d)
- 2
Description
## What happens
Background subagents driving long tasks for this repo (e2e suites, `record` invocations, PR triage) repeatedly stop mid-task with a final message like *"the e2e run will re-invoke me on completion"* / *"awaiting the monitor's completion event"* — and the promised wake-up never comes. The coordinator receives a completed-notification while the OS process (the e2e suite, the recording) is still running. The work sits half-done for 10–40 minutes until a human or the coordinator resumes the agent, which then finds its command finished long ago.
**Observed six times over 2026-07-24 → 2026-07-27**, across unrelated agents (a dogfood run, a PR-feedback triage, a PR-merge task, and the headless-migration implementation — the last stalled twice, once straight through an explicit "poll directly, don't wait" instruction).
## Root cause (investigated, mechanism confirmed on Claude Code 2.1.219)
The background-task re-invocation contract is honored only for a session that is **still live**:
- The **top-level session** idles between user turns, so it survives ending a turn with a `run_in_background` Bash task pending, and is re-invoked when that task exits.
- A **subagent has no idle state.** The moment it emits a final message with no in-flight foreground tool call, `SubagentStop` fires and the subagent is finalized — its result returns to the parent, and its still-running background task is **orphaned**. A background Bash task does not count as a "live background child" that defers finalization; only an in-flight child *agent* would, and background subagents have the Agent tool filtered out, so they can't create one.
- **Positive control (the load-bearing measurement):** the completion notification *is* delivered to a subagent that stays live and keeps working — verified directly with a 90s background task whose completion notification arrived at the next tool-step boundary. The wiring exists; turn-end finalization races ahead of it.
So every stall is the same one-line mistake: *the agent ended its turn expecting to be woken, and turn-end is precisely the thing that makes waking impossible.*
## Mitigations in place
1. **Brief discipline (works when followed):** long commands run as ONE foreground Bash call with an explicit timeout (a ~90s e2e suite fits comfortably in `timeout: 600000`), or are driven to completion inside a single Bash call with a bounded poll:
```bash
for attempt in $(seq 1 240); do [ -s out.txt ] && break; sleep 5; done; cat out.txt
```
(`-s` so a half-written file isn't read; foreground `sleep` inside one call is allowed for subagents — verified.)
2. **Mechanical guard (built and tested, NOT yet installed — see TODO):** a `SubagentStop` hook in the main session's project settings that blocks a stall-shaped final message (`re-invoke` / `will notify` / `awaiting … completion` …) and sends back the poll instruction as the agent's next input. Loop-guarded via `stop_hook_active` (one nudge, then allow). Script logic fully tested against mock hook payloads; the hook runs in the **main session**, so it does not depend on subagents inheriting project hooks.
## Human TODO
- [ ] **Install the stall guard** (agent self-installation of hook config is classifier-blocked, correctly):
1. Copy `subagent-stall-guard.sh` into `.claude/hooks/` and `chmod +x` it (the tested script is in the session scratchpad under `stall-investigation/sandbox-project/.claude/hooks/`; it is ~40 lines, pure bash, no dependencies — reproduced below for permanence).
2. Add to `.claude/settings.local.json`:
```json
"hooks": { "SubagentStop": [ { "matcher": "*", "hooks": [ { "type": "command", "command": "$CLAUDE_PROJECT_DIR/.claude/hooks/subagent-stall-guard.sh" } ] } ] }
```
- [ ] **File the upstream issue** on the Claude Code repository (draft below — the harness asymmetry is theirs to fix properly).
- [ ] Optional: decide whether the hook should be committed for all contributors (`.claude/settings.json`) once it has proven itself locally.
## The hook script (for permanence)
```bash
#!/usr/bin/env bash
# SubagentStop guard: catch a subagent that ends its turn expecting to be
# re-invoked when a background task finishes, and force it to poll instead.
set -euo pipefail
input="$(cat)"
if printf '%s' "$input" | grep -qE '"stop_hook_active"[[:space:]]*:[[:space:]]*true'; then
exit 0
fi
if printf '%s' "$input" | grep -qiE 're-?invoke|will notify|notify (me )?(on|upon) completion|awa(it|iting)|waiting (for|on) .*(background|task|run|monitor|completion|event)|monitor.?s? (completion|event)|(both|the).*(will )?notify'; then
printf '%s\n' '{"decision":"block","reason":"A background task looks like it is still running and you ended your turn expecting to be re-invoked on its completion. Subagents are NOT re-invoked when a run_in_background Bash task finishes: you are finalized on turn-end and the task is orphaned. Do not wait. In ONE Bash tool call, poll the output file to completion, then act on it this turn, e.g.: for attempt in $(seq 1 240); do [ -s OUTFILE ] && break; sleep 5; done; cat OUTFILE"}'
exit 0
fi
exit 0
```
## Upstream issue draft (ready to file)
> **Title:** A subagent with a live background Bash task is finalized on turn-end instead of re-invoked on task completion
>
> **Version:** 2.1.219
>
> **Summary:** A top-level session that starts a `run_in_background` Bash task and ends its turn is re-invoked when the task completes. A subagent doing the same is finalized: `SubagentStop` fires on turn-end, the result returns to the parent, and the still-running background task is orphaned. The completion notification never re-invokes the subagent.
>
> **Repro:** (1) spawn a background subagent whose brief says to run a long command with `run_in_background: true` and then end its turn "to be re-invoked on completion"; (2) the coordinator receives the subagent-completion notification while the OS process demonstrably still runs; (3) resuming the subagent shows `stopped (completed)`; it finds the command finished long ago and only then completes the work.
>
> **Expected:** either hold the subagent pending like the top-level session and re-invoke it on task completion, or refuse to finalize a subagent with a live tracked background task.
>
> **Actual:** finalized on turn-end; task orphaned; `&`-spawned "waiters" reaped only on a later resume.
>
> **Confirmed sub-facts:** the completion notification IS delivered to a subagent that stays live (verified with a 90s background task); background subagents have the Agent tool filtered, so a foreground poll loop inside one tool call is their only wait primitive.
Contributor guide
Research direction
Start with the reproduced subagent-stall-guard.sh script and the .claude/settings.local.json hook configuration described in the issue. Install the executable under .claude/hooks/, verify the guard against the documented mock payloads and a stalled background-task message, and confirm that normal stops still pass through; the upstream issue draft is ready to file separately.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- bash
- Domain
- developer-experience, tooling
- Issue type
- Bug
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 55/100