github / github/copilot-cli

Session resume triggers background compaction that fails silently and hangs the process indefinitely (recurred 4x)

未关闭
#4,138 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area:context-memory area:sessions
主要语言
Shell
星标
11.2k
派生
1.9k
平均合并
14 小时 16 分钟
30 天内合并 PR
6

描述

### Describe the bug

On session **resume** (not manual `/compact`), the CLI's background `CompactionProcessor` immediately attempts to compact the pre-existing conversation history. If that compaction call to the model returns an empty response, there is no retry or graceful fallback — the process hangs **indefinitely** with no further output, and must be force-killed (closing stdin) to recover. This has recurred **4 times across different sessions** for me.

This differs from #2500 (session stops and asks the user to repeat their request) and #2861 (three failures on manual `/compact`, but the session remains usable and a later `/compact` succeeds) — in my case the process never recovers and never becomes responsive again.

### Affected version

1.0.71-2

### Steps to reproduce the behavior

1. Have a session/chat with meaningful prior conversation history.
2. Resume that session (e.g. reopening a chat in the desktop app, or `copilot --resume`).
3. MCP servers reconnect (in my case ~9-10 MCPs, including one - Azure MCP via `npx -y @azure/mcp@latest` - that stalled 39s due to a Windows `EPERM: operation not permitted, unlink ...azmcp.exe` file-lock error from a stale/lingering `azmcp.exe` process).
4. Background `CompactionProcessor` kicks off automatically against the resumed history.
5. Observe two consecutive `Compaction failed: received empty response from model` entries in the process log.
6. The process log goes silent afterward (no more entries) even though the process is technically still running - it never responds to input again.

Log excerpt (`process--.log`):
```
15:09:15 session resume begins, MCPs reconnecting
15:09:56 Azure MCP connects (39127ms - EPERM unlink stall on azmcp.exe)
15:10:03 first completion request post-resume
15:10:10 CompactionProcessor: Compaction failed: received empty response from model
15:10:17 CompactionProcessor: Compaction failed: received empty response from model
(silence - process never recovers)
15:18:43 manual shutdown (stdin closed)
```

### Expected behavior

Background compaction failures on resume should not hang the whole session. At minimum:
- Retry with backoff, or fall back to serving the resumed session without a fresh compacted summary.
- Surface the failure to the user/UI instead of silently hanging.
- Ensure the process remains responsive to new input even if background compaction never succeeds.

### Additional context

- OS: Windows
- Slow/stalled MCP connections (e.g. Azure MCP taking 30-40s to connect due to a stale `azmcp.exe` process holding a file lock) may be a contributing/confounding timing factor around when the first post-resume compaction request fires, but the core defect is the lack of retry/fallback/timeout on `CompactionProcessor` failure.
- Ruled out as causes: total registered MCP tool schema count (~400+, shared with a healthy long-running session that never crashed) and accumulated tool-call-result payload size (the process log was only 26.2KB/142 lines at time of failure).
- Related: #2500 (compaction on resume causing session confusion), #2861 (manual `/compact` triple-failure, but recoverable).

贡献指南

打开贡献指南

调研方向

从会话恢复期间触发的 CompactionProcessor 路径开始,检查如何处理空的模型响应。在观察进程日志的同时使用恢复的会话重现问题,然后验证失败的后台压缩不会使会话失去响应,并且会报告或绕过该失败。

由索引模型根据 Issue 内容生成。

评估

技术栈
shell
领域
cli
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
冷清
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。