CommandCodeAI / CommandCodeAI/command-code
OOM crash (heap out of memory) when resuming long session [1.39.2]
还没有人认领这个 Issue。
- 主要语言
- 没有语言数据
- 星标
- 4k
- 派生
- 350
- PR 合并指标
- 30 天内没有已合并 PR
描述
Description
Running cmd --resume <session-id> --yolo on a long-running agent session crashes with a Node.js FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. The process aborts (zsh: abort) and the session is lost.
Environment
- Command Code: 1.39.2
- Node: v24.19.0 (
/Users/shreyashphakadepawar/.nvm/versions/node/v24.19.0/bin/node) - OS: macOS 26.6.2 (Build 25G83), arm64
- Invocation:
cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo - Default Node heap limit on this machine: 2240 MB (
v8.getHeapStatistics().heap_size_limit)
Crash output
<--- Last few GCs --->
[20041:0x75280c000] 68684349 ms: Mark-Compact 1938.6 (2098.2) -> 1923.8 (2098.0) MB, pooled: 1 MB, 31.92 / 0.04 ms (average mu = 0.305, current mu = 0.306) task; scavenge might not succeed
[20041:0x75280c000] 68684395 ms: Mark-Compact 1939.8 (2098.0) -> 1922.8 (2097.1) MB, pooled: 1 MB, 32.96 / 0.04 ms (average mu = 0.300, current mu = 0.296) allocation failure; scavenge might not succeed
FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
----- Native stack trace -----
1: 0x1043219f4 node::OOMErrorHandler(char const*, v8::OOMDetails const&) [...]
...
39: 0x188ef84e4 start [/usr/lib/dyld]
zsh: abort cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
What happened / steps to reproduce
- Run a long-lived agent conversation (this one had been active across compaction, with a large session transcript and many tool calls — screenshots, sub-agents, image reads).
- Resume that session headlessly:
cmd --resume <session-id> --yolo. - The process climbs to ~1.94 GB, GC runs become ineffective (mark-compact can't reclaim), and Node aborts with OOM.
Suspected cause
The resume path likely loads the full session transcript (which after many turns / compaction / large tool outputs — including read_file image data and agent results — grows large) into memory after the shared/global heap is already near the default Node limit. Because the process was launched by node without --max-old-space-size, it dies at the default 2 GB-ish ceiling rather than degrading gracefully. Possible contributors:
- Session transcript/context payload loaded in full on
--resume. - Tool outputs (screenshots/images/large read results from sub-agents) retained in memory instead of being streamed or evicted.
- No explicit heap sizing or soft-cap/memory-pressure handling in the resume path.
Expected behavior
A long session resume should either:
- not exceed the default heap (stream/evict large payloads), or
- surface a clear, recoverable error (e.g. "session too large, please compact/start fresh") instead of a hard
abortthat loses the work, or - raise/self-tune the heap limit (e.g.
--max-old-space-size) or chunk the resume.
Additional context
This happened on a session that had already been auto-compacted once and contained a large number of tool calls including image reads. It is the first reproducible-ish OOM in ordinary usage — no custom memory settings were in place.
Impact
- The resumed session crashes and the user loses the working context.
- Error message only points at the JS engine, not at Command Code — no hint to reduce session size or compact first.
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 cmd --resume <session-id> --yolo 恢复路径开始,并使用一个包含压缩、大型工具输出和图像读取的长会话进行复现。跟踪恢复期间会话 transcript 和工具结果是如何加载和保留的。完成的标准是会话避免 default heap OOM,或者报告清晰且可恢复的错误,而不是中止并丢失会话。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- javascript, node.js
- 领域
- cli, performance
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 48/100