Session resume triggers background compaction that fails silently and hangs the process indefinitely (recurred 4x)
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Mô tả
### Describe the bug
On session **resume** (not manual `/compact`), the CLI's background `CompactionProcessor` immediately attempts to compact the pre-existing conversation history. If that compaction call to the model returns an empty response, there is no retry or graceful fallback — the process hangs **indefinitely** with no further output, and must be force-killed (closing stdin) to recover. This has recurred **4 times across different sessions** for me.
This differs from #2500 (session stops and asks the user to repeat their request) and #2861 (three failures on manual `/compact`, but the session remains usable and a later `/compact` succeeds) — in my case the process never recovers and never becomes responsive again.
### Affected version
1.0.71-2
### Steps to reproduce the behavior
1. Have a session/chat with meaningful prior conversation history.
2. Resume that session (e.g. reopening a chat in the desktop app, or `copilot --resume`).
3. MCP servers reconnect (in my case ~9-10 MCPs, including one - Azure MCP via `npx -y @azure/mcp@latest` - that stalled 39s due to a Windows `EPERM: operation not permitted, unlink ...azmcp.exe` file-lock error from a stale/lingering `azmcp.exe` process).
4. Background `CompactionProcessor` kicks off automatically against the resumed history.
5. Observe two consecutive `Compaction failed: received empty response from model` entries in the process log.
6. The process log goes silent afterward (no more entries) even though the process is technically still running - it never responds to input again.
Log excerpt (`process--.log`):
```
15:09:15 session resume begins, MCPs reconnecting
15:09:56 Azure MCP connects (39127ms - EPERM unlink stall on azmcp.exe)
15:10:03 first completion request post-resume
15:10:10 CompactionProcessor: Compaction failed: received empty response from model
15:10:17 CompactionProcessor: Compaction failed: received empty response from model
(silence - process never recovers)
15:18:43 manual shutdown (stdin closed)
```
### Expected behavior
Background compaction failures on resume should not hang the whole session. At minimum:
- Retry with backoff, or fall back to serving the resumed session without a fresh compacted summary.
- Surface the failure to the user/UI instead of silently hanging.
- Ensure the process remains responsive to new input even if background compaction never succeeds.
### Additional context
- OS: Windows
- Slow/stalled MCP connections (e.g. Azure MCP taking 30-40s to connect due to a stale `azmcp.exe` process holding a file lock) may be a contributing/confounding timing factor around when the first post-resume compaction request fires, but the core defect is the lack of retry/fallback/timeout on `CompactionProcessor` failure.
- Ruled out as causes: total registered MCP tool schema count (~400+, shared with a healthy long-running session that never crashed) and accumulated tool-call-result payload size (the process log was only 26.2KB/142 lines at time of failure).
- Related: #2500 (compaction on resume causing session confusion), #2861 (manual `/compact` triple-failure, but recoverable).
Hướng dẫn đóng góp
Hướng nghiên cứu
Bắt đầu từ đường dẫn CompactionProcessor được kích hoạt trong quá trình tiếp tục phiên và kiểm tra cách các phản hồi rỗng của model được xử lý. Tái hiện bằng một phiên được tiếp tục trong khi theo dõi process log, sau đó xác minh rằng một lần compact background thất bại vẫn để phiên hoạt động bình thường và báo cáo hoặc bỏ qua lỗi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- shell
- Lĩnh vực
- cli
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Ít trao đổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 48/100