anthropics / anthropics/claude-code
No stream-inactivity watchdog: unattended sessions hang forever on wedged API response; headless --print dies terminally on retryable 429/5xx
- Ngôn ngữ chính
- Python
- Star
- 145k
- Fork
- 23.1k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
## Summary
Unattended Claude Code sessions on our host repeatedly go permanently silent mid-task. We have now root-caused three operator-visible "stream halt" incidents to two client resilience gaps around API stream failures. Local workarounds have been built three times; the failure class recurs because both gaps are client-side.
## Environment
- Claude Code 2.1.234 (CLI + VS Code extension 2.1.234), Linux 6.8.0-117-generic
- Models: claude-fable-5 (interactive), claude-opus-5 (headless subprocesses)
- Usage pattern: long-running interactive tmux sessions + headless `claude --print --output-format stream-json` worker subprocesses
## Defect 1: no inactivity watchdog on an in-flight API stream (interactive)
An interactive session's API response stream wedged mid-response and the client spun on the "Whisking…" spinner for **26+ minutes** with no timeout, no retry, and no surfaced error. The session was unattended (tmux); to an operator the stream is simply silent forever. Killing/resuming the session is the only recovery.
Observed today (2026-08-18 ~15:59–16:26 UTC): spinner showed `Whisking… (26m 11s · ↓ 60.2k tokens)` with zero child processes and no network progress; the turn never completed.
Expected: a stream inactivity timeout (e.g. no bytes for N seconds) that aborts, surfaces a typed error, and retries/resumes — especially critical for unattended/headless operation.
## Defect 2: headless `--print` runs terminate fatally on retryable errors
Headless stream-json runs die terminally on errors that should be retryable, discarding long in-flight tasks:
- `"result":"API Error: Server error mid-response. The response above may be incomplete."`, `"terminal_reason":"api_error"`, `"api_error_status":null` — after **50 turns, 698s of API time, $4.93 spend**, the run simply ends (and confusingly reports `"subtype":"success"` in the result JSON).
- Three more runs the same day died with `"api_error_status":429`, `"terminal_reason":"api_error"` — a rate-limit, i.e. the textbook retryable error, treated as terminal mid-task.
Expected: retry with backoff on 429/5xx mid-run (or a resume mechanism for `--print` runs), and `subtype` should not be `success` when `terminal_reason` is `api_error`.
## Impact
- 9 headless worker runs across 2026-08-17/18 ended with `terminal_reason:"api_error"`, several after multi-dollar spend, requiring full re-dispatch.
- Interactive unattended sessions stall silently 4–7 turns in; a human has to notice and take over. This has happened on 3 consecutive days despite local mitigations (liveness sweeps, heartbeat/first-contact logging, watchdog scripts).
## Repro sketch
Long-running unattended session or `--print` run on a busy account (concurrent subprocess workers). Any mid-response 5xx wedge or 429 reproduces the class; the client behavior (indefinite spinner / terminal death) is deterministic once the stream fails.
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
Start by tracing the headless `claude --print --output-format stream-json` flow and the interactive API response stream handling. Reproduce a mid-response 429/5xx and an inactive stream, then identify the existing error and retry paths. Done means inactivity is surfaced and recoverable, retryable errors resume with backoff, and `api_error` results are not reported as success.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python, vscode
- Lĩnh vực
- api, cli
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 48/100