Session resume triggers background compaction that fails silently and hangs the process indefinitely (recurred 4x)
- 主要言語
- Shell
- スター
- 11.2k
- フォーク
- 1.9k
- 平均マージ
- 14時間 16分
- マージ済み PR(30日)
- 6
説明
### Describe the bug
On session **resume** (not manual `/compact`), the CLI's background `CompactionProcessor` immediately attempts to compact the pre-existing conversation history. If that compaction call to the model returns an empty response, there is no retry or graceful fallback — the process hangs **indefinitely** with no further output, and must be force-killed (closing stdin) to recover. This has recurred **4 times across different sessions** for me.
This differs from #2500 (session stops and asks the user to repeat their request) and #2861 (three failures on manual `/compact`, but the session remains usable and a later `/compact` succeeds) — in my case the process never recovers and never becomes responsive again.
### Affected version
1.0.71-2
### Steps to reproduce the behavior
1. Have a session/chat with meaningful prior conversation history.
2. Resume that session (e.g. reopening a chat in the desktop app, or `copilot --resume`).
3. MCP servers reconnect (in my case ~9-10 MCPs, including one - Azure MCP via `npx -y @azure/mcp@latest` - that stalled 39s due to a Windows `EPERM: operation not permitted, unlink ...azmcp.exe` file-lock error from a stale/lingering `azmcp.exe` process).
4. Background `CompactionProcessor` kicks off automatically against the resumed history.
5. Observe two consecutive `Compaction failed: received empty response from model` entries in the process log.
6. The process log goes silent afterward (no more entries) even though the process is technically still running - it never responds to input again.
Log excerpt (`process--.log`):
```
15:09:15 session resume begins, MCPs reconnecting
15:09:56 Azure MCP connects (39127ms - EPERM unlink stall on azmcp.exe)
15:10:03 first completion request post-resume
15:10:10 CompactionProcessor: Compaction failed: received empty response from model
15:10:17 CompactionProcessor: Compaction failed: received empty response from model
(silence - process never recovers)
15:18:43 manual shutdown (stdin closed)
```
### Expected behavior
Background compaction failures on resume should not hang the whole session. At minimum:
- Retry with backoff, or fall back to serving the resumed session without a fresh compacted summary.
- Surface the failure to the user/UI instead of silently hanging.
- Ensure the process remains responsive to new input even if background compaction never succeeds.
### Additional context
- OS: Windows
- Slow/stalled MCP connections (e.g. Azure MCP taking 30-40s to connect due to a stale `azmcp.exe` process holding a file lock) may be a contributing/confounding timing factor around when the first post-resume compaction request fires, but the core defect is the lack of retry/fallback/timeout on `CompactionProcessor` failure.
- Ruled out as causes: total registered MCP tool schema count (~400+, shared with a healthy long-running session that never crashed) and accumulated tool-call-result payload size (the process log was only 26.2KB/142 lines at time of failure).
- Related: #2500 (compaction on resume causing session confusion), #2861 (manual `/compact` triple-failure, but recoverable).
コントリビューションガイド
調査の方向性
セッション再開中にトリガーされる CompactionProcessor のパスから調査を始め、空のモデル応答がどのように処理されるかを確認します。プロセスログを監視しながら再開したセッションで再現し、その後、バックグラウンドのコンパクションが失敗してもセッションが応答可能な状態を保ち、失敗を報告するか回避することを確認します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- shell
- 領域
- cli
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 48/100