anthropics / anthropics/claude-code

Background daemon: fresh-spawned (non-spare) --bg workers crash "exit N before init" 100% of the time

未关闭
#90,149 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area:agent-view bug has repro platform:macos
主要语言
Python
星标
145k
派生
23.1k
PR 合并指标
PR 指标待抓取

描述

**Version:** 2.1.212 (also observed on 2.1.14x-2.1.15x daemons; macOS 14, Darwin 23.6.0)

**Summary**

The background daemon keeps one pre-warmed spare worker. Jobs that claim the spare (`bg claimed-spare ` in daemon.log) succeed. Jobs dispatched through the fresh-spawn path (`bg spawned `) crash roughly 13 seconds later with `bg settled (crashed): exit N before init`, before session init, every single time. In my full daemon.log history: 394 spare-claimed jobs, almost all successful; 16 fresh-spawned jobs, 16 crashed. Zero fresh-spawn successes on record.

**Reproduction conditions**

A job lands on the fresh-spawn path in two cases, both reproducible on my machine:

1. Two `claude --bg` dispatches arrive in the same second. The first claims the spare and succeeds; the second fresh-spawns and crashes. Observed daily when two scheduled jobs fire at the same minute; the crashing job alternates depending on which dispatch wins the race.
2. Cold daemon start (no spare exists yet). The dispatching job fresh-spawns and crashes.

**Sample log window (2026-08-27, UTC)**

```
14:00:05.726 [bg] bg claimed-spare d8fe41c4 (shell) <- succeeded (ran 7m, done)
14:00:05.745 [bg] bg spawned f6eba5ef (shell) <- fresh spawn
14:00:05.760 [bg] bg spare spawned host pid=16857
14:00:18.794 [bg] bg settled f6eba5ef (crashed): exit 1 before init
```

Auth was healthy in the same window (`auth: scheduling proactive refresh in 28560s`, no failure lines), so this is not the headless OAuth-refresh issue. Immediately after past occurrences, direct `claude -p` and `claude --bg` probes with the identical flag set completed clean, so flags, model, prompt, and credentials are not the cause.

**Impact**

Any consumer that schedules more than one background job at the same minute, or dispatches into a cold daemon, deterministically loses a job. The worker's stderr is not observable from the dispatching side (the launcher's pipes are destroyed at the `backgrounded` banner), so the only diagnostic available is the `exit N before init` detail string.

**Ask**

Either fix the fresh-spawn worker path, or have the daemon queue the second dispatch until the replenishment spare is ready, or surface the worker's early stderr in `state.json` so consumers can see why init failed.

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start by reproducing the cold-daemon and two-dispatch cases with `claude --bg`, then inspect the daemon log entries for `bg spawned`, `bg claimed-spare`, and `exit N before init`. Trace the fresh-spawn path and the `state.json` reporting path; done means fresh-spawned workers initialize reliably or their early failure is made observable, with queued dispatches handled if that is the chosen behavior.

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
backend, cli
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
活跃
描述清晰度
需要澄清
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。