Terminal UI stops consuming events (input + scroll dead) when a turn spawns parallel subagents; runtime keeps running
还没有人认领这个 Issue。
- 主要语言
- Shell
- 星标
- 11.2k
- 派生
- 1.9k
- 平均合并
- 14 小时 16 分钟
- 30 天内合并 PR
- 6
描述
Summary
On the prerelease channel (1.0.81-4, 1.0.81-5), the terminal UI stops
consuming runtime events at the moment a turn launches a parallel block of
subagents. The Rust runtime is unaffected and keeps working — subagents
continue making model calls for minutes afterwards and their results are
persisted to the session store — but the terminal never repaints again, accepts
no keyboard input, and scrollback is frozen.
The process is left idle, not spinning: 0 CPU ticks over a 5s sample,
State: S (sleeping), wchan: futex_wait_queue_me, with no busy child
processes. It never recovers and does not exit on its own.
Reproduced 3 times in one evening across independent sessions.
Environment
- Copilot CLI
1.0.81-5(also seen on1.0.81-4) — prerelease channel - Node.js v22.21.1, Linux x64
- Several MCP servers configured (stdio + http)
Steps to reproduce
- Run a workflow where the top-level agent invokes a custom agent via
task. - Have that custom agent fan out to 2–3 subagents of its own in a single
parallel block (i.e. multipletaskcalls issued in one turn). - Observe the terminal at the moment the parallel block is launched.
Total concurrency at failure was 3–4 simultaneous subagents. Sequential
subagent execution has not reproduced it.
Evidence
1. The UI log stops at a turn boundary, always with the same event triple.
The final lines in every frozen session:
[WARNING] useTimeline: skipping unprocessable event type=model.response: Error: Unhandled event type in timeline: model.response
[WARNING] useTimeline: skipping unprocessable event type=model.turn_ended: Error: Unhandled event type in timeline: model.turn_ended
[WARNING] useTimeline: skipping unprocessable event type=model.messages_snapshot: Error: Unhandled event type in timeline: model.messages_snapshot
2. The timeline reducer cannot handle any model.* event. These warnings
are continuous throughout normal operation, thousands per session — one per
runtime event. Counts scale directly with subagent usage (0–260 in ordinary
sessions; 2,504 in a heavy multi-agent session). This looks like event types
emitted by the runtime that the TUI timeline reducer does not know about, i.e.
a skew between the runtime and the UI layer within the same build.
Observed types: model.message, model.turn_started, model.turn_ended,
model.model_call_started, model.model_call_success,
model.captured_assignment_context, model.tool_execution, model.response,
model.messages_snapshot.
3. The runtime continues after the UI dies. Per-call usage rows recorded in
the local session store, cross-referenced with the last UI log line:
| Session | Last UI event | Runtime kept working until | Peak concurrent subagents |
|---|---|---|---|
| A | T+0 | — | 3 (reached that same minute) |
| B | T+0 | T+8min | 4 |
| C | T+0 | T+2min or more | 4 (fan-out began T+1min) |
In session B the subagents ran to completion and their results were persisted;
only the UI was lost. copilot --resume <session-id> recovers the session,
which confirms the runtime state is intact.
Expected
The UI keeps rendering and accepting input while subagents run in parallel, or
at minimum fails loudly rather than silently detaching from the event stream.
Actual
UI silently stops consuming events. Terminal appears hung; the only recovery is
killing the process and resuming the session.
Secondary issue: frozen sessions ignore SIGTERM and can leave a spinning zombie
- All four affected processes ignored
SIGTERMand requiredSIGKILL. - One process that had begun disposal never exited, logging this once per
second indefinitely (grew its log to 656 KB before it was killed):
[ERROR] SessionClient poll loop error (will retry): Error: Cannot invoke native session after disposal has started
[ERROR] SessionClient background task refresh failed (continuing) [phase=pre-dispatch cursor=… attempts=2
Over one session these accumulated to 10 simultaneous long-lived CLI processes,
none of which could be shut down normally.
Workaround
Force subagents to run sequentially rather than in parallel blocks, and pin to
the stable channel (COPILOT_AUTO_UPDATE=false with a stable version installed
via npm).
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
首先复现并行 custom-agent 工作流,并跟踪所列 model.* 事件周围的 TUI 时间线 reducer 和事件流处理。将 reducer 的处理方式与日志中观察到的运行时事件进行比较;完成的标准是:在并行 subagents 运行期间,UI 仍能持续渲染并接受输入,且受影响的进程能够正常关闭。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- node.js, rust, shell
- 领域
- backend, cli
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100