OOM crash (`JavaScript heap out of memory`) on long `--resume` sessions; crash dumps written into the user's cwd
还没有人认领这个 Issue。
- 主要语言
- Shell
- 星标
- 11.2k
- 派生
- 1.9k
- 平均合并
- 14 小时 16 分钟
- 30 天内合并 PR
- 6
描述
Describe the bug
Summary
Copilot CLI 1.0.82 repeatedly dies with a V8 heap OOM during long resumed sessions. It crashed 3 times in ~14 hours for me, always at the 4 GiB heap cap.
Separately, the resulting Node diagnostic reports are written into the current working directory, so they land inside whatever git repo I'm working in and show up as untracked files.
Environment
| Copilot CLI | 1.0.82 |
| Node | v24.18.1 |
| Platform | linux x64 (glibc 2.39) |
| Invocation | copilot --no-warnings --report-on-fatalerror --optimize-for-size --expose-gc copilot --resume |
What happens
"event": "Allocation failed - JavaScript heap out of memory"
"trigger": "OOMError"
Three crashes on 2 Sep 2026, all in resumed sessions:
| Time (UTC) | Heap used / limit | RSS | CPU (user s) |
|---|---|---|---|
| 00:51:37 | 3.98 / 4.00 GiB | 4.63 GiB | 7833 |
| 03:08:30 | 3.99 / 4.00 GiB | 4.75 GiB | 4065 |
| 14:39:00 | 3.99 / 4.00 GiB | 4.69 GiB | 2785 |
Diagnostic detail
At crash, old_space alone accounted for 3.96 GiB of the 3.99 GiB used:
old_space: 3.96 GiB
code_space: 0.01 GiB
trusted_space: 0.01 GiB
large_object_space: 0.00 GiB
Nearly all of it is long-lived, survived-GC data — consistent with unbounded retention/a leak rather than a transient allocation spike. javascriptStack is "No stack." / Unavailable., as expected for an OOM abort.
All three crashes were --resume sessions with substantial CPU time behind them (46m–2h11m user CPU), which points at per-session state (conversation/tool history?) accumulating without bound. Heap is pinned right at the cap each time, and RSS exceeds it by ~0.6–0.75 GiB.
Impact
- Data/work loss — long sessions die abruptly once they get big enough. The practical ceiling appears to be session length, not task complexity.
- Repo pollution — because
--report-on-fatalerroris set and Node writes reports relative to cwd, files likereport.20260902.143900.310882.0.001.jsonappear inside the user's repository. They show as untracked ingit statusand are easy to commit by accident. A CLI shouldn't drop crash artifacts into arbitrary project directories.
Notes for Copilot debugging
The 4 GiB ceiling is not user-imposed:
NODE_OPTIONSis unset.- Host has 30 GiB RAM with ~9 GiB free at crash time — no system memory pressure.
- Plain
nodeon this host reports a 4.05 GiB default heap limit;node --optimize-for-sizereports exactly 4.00 GiB, matchingjavascriptHeap.memoryLimitin every dump.
So the cap comes from the --optimize-for-size flag the CLI passes itself. The process aborts at a self-imposed ceiling while the machine still has memory to spare — worth considering whether that flag is right for long-lived sessions, though the underlying retention growth looks like the real bug.
Affected version
1.0.82
Steps to reproduce the behavior
- Start a Copilot CLI session in a git repo.
- Use it heavily / resume it over an extended period (mine ran for hours of CPU time).
- Session eventually aborts with the OOM above; a
report.*.jsonappears in the repo root.
Expected behavior
- Session memory should be bounded (evict/compact old history), or degrade gracefully instead of a hard abort.
- Crash dumps should go somewhere tool-owned — e.g.
~/.copilot/logs/or$TMPDIR— not the user's cwd. If cwd is intentional, it should be configurable.
Additional context
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
首先,使用 --report-on-fatalerror 和 --optimize-for-size 重现一个较长的 --resume 会话,然后检查随附的 Node 诊断报告和会话历史记录行为。完成的标准是:长会话不再持续增长直到发生不可恢复的 OOM,并且致命错误报告会写入工具拥有的位置,而不是用户的 cwd。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- javascript, nodejs
- 领域
- cli, performance
- Issue 类型
- 缺陷
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 35/100