v1.0.80: event-storage exhaustion retry storm drives long-running session into GC/compaction loop and Node OOM
还没有人认领这个 Issue。
- 主要语言
- Shell
- 星标
- 11.2k
- 派生
- 1.9k
- 平均合并
- 14 小时 16 分钟
- 30 天内合并 PR
- 6
描述
Describe the bug
A long-running active Copilot CLI session eventually reached remote event-storage exhaustion. After that, the exporter kept attempting 500-event flushes, while the process repeatedly reported memory pressure, forced GC/emergency compaction, and thousands of bridge acknowledgement timeouts. The CLI ultimately terminated with a Node out-of-memory error shown in the terminal.
This appears to connect two existing failure modes that are currently reported separately:
- #4467:
session_event_storage_exhausted - #4506: memory-pressure GC/compaction loop ending in OOM
- Also related to #4251 for unbounded memory growth in large sessions
The new observation is that remote storage exhaustion does not merely make session status unreliable: the repeated exporter failures/backlog correlate with sustained local memory pressure and eventual process death.
Environment
- Copilot CLI at failure: 1.0.80
- Embedded Node.js: v24.18.1
- OS: Linux x86_64, kernel 6.17.0-1022-azure
- Host memory: 125 GiB RAM, no swap
- Session created: 2026-08-10
- Process/resume started: 2026-08-18
- Failure: 2026-08-27 (about 8.8 days of this process, about 16.9 days of session lifetime)
- Workload: long-running interactive/autonomous software-engineering session with many tool calls and subagents
Session size at failure
events.jsonl: 934,639,321 bytes (~892 MiB)- Event records: 282,912
- Checkpoints: 32
Log evidence
Across the process log:
| Event | Count |
|---|---|
session_event_storage_exhausted |
18,909 |
Memory pressure detected - requesting garbage collection |
21,352 |
Memory pressure persists after GC - triggering emergency compaction |
473 |
timed out waiting for bridge event ack |
3,001 |
Representative final minutes:
[WARNING] Failed to submit events to Mission Control session …: 409
{"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
[WARNING] remote session batch flush failed {"event_count":500,...}
[WARNING] remote session exporter circuit opened after batch flush failure
[WARNING] Memory pressure detected - requesting garbage collection
[WARNING] Memory pressure persists after GC - triggering emergency compaction
[WARNING] timed out waiting for bridge event ack {"timeout_ms":30000}
[INFO] Compacted 17 messages, saved ~9267 tokens
[WARNING] Memory pressure detected - requesting garbage collection
...
[INFO] Compacted 38 messages, saved ~23339 tokens
[WARNING] Memory pressure detected - requesting garbage collection
The application log ends in repeated memory-pressure warnings. The fatal Node OOM text was printed by the terminating runtime to the terminal and was not persisted into the Copilot process log. No kernel OOM-killer event was recorded, so this was a Node/V8 process failure rather than the Linux kernel killing the process.
The host currently has ample free memory after restart; I did not capture /proc/<pid>/status immediately before the crash, so I cannot claim whether the final allocation failure was at the V8 heap limit or in native/external memory.
Steps to reproduce
- Keep one Copilot CLI session active/resumable for more than a week with frequent tool calls and subagent events.
- Allow
events.jsonlto grow toward 1 GB and remote Mission Control event storage to fill. - Observe repeated 409
session_event_storage_exhaustedresponses for 500-event batches. - Continue using the session.
- Observe repeated memory-pressure GC, emergency compaction, bridge ack timeouts, and eventually a Node OOM termination.
Expected behavior
- Rotate/continue/compact remote event storage before its limit is reached.
- Once the server permanently rejects a stream as exhausted, stop retaining/retrying the rejected backlog in a way that grows local memory.
- Keep local session operation and checkpointing healthy when remote export is unavailable.
- Apply bounded backoff and expose a clear degraded-export state.
- Do not repeatedly perform lossy conversation compaction unless it measurably relieves the memory pressure.
- Persist the fatal Node diagnostic report path or equivalent crash telemetry in the process log.
Actual behavior
A permanent remote quota condition produced 18,909 failed submissions, concurrent memory-pressure/compaction loops, degraded bridge responsiveness, and eventual Node OOM. Recovery required starting a new session and manually reading the old checkpoint and session directory.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从远程会话导出器的500事件flush失败、session_event_storage_exhausted响应以及相关的process-log内存压力消息开始。跟踪被拒绝的批次如何被保留和重试,然后验证永久性导出失败不会影响本地会话操作的健康状态,使用有界重试行为,并记录清晰的降级状态,同时不会导致OOM。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- node.js
- 领域
- backend, cli, performance
- Issue 类型
- 缺陷
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 28/100