CommandCodeAI / CommandCodeAI/command-code
OOM crash (heap out of memory) when resuming long session [1.39.2]
還沒有人認領這個 Issue。
- 主要語言
- 沒有語言資料
- 星號
- 4k
- 分支
- 350
- PR 合併指標
- 30 天內沒有已合併 PR
描述
Description
Running cmd --resume <session-id> --yolo on a long-running agent session crashes with a Node.js FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. The process aborts (zsh: abort) and the session is lost.
Environment
- Command Code: 1.39.2
- Node: v24.19.0 (
/Users/shreyashphakadepawar/.nvm/versions/node/v24.19.0/bin/node) - OS: macOS 26.6.2 (Build 25G83), arm64
- Invocation:
cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo - Default Node heap limit on this machine: 2240 MB (
v8.getHeapStatistics().heap_size_limit)
Crash output
<--- Last few GCs --->
[20041:0x75280c000] 68684349 ms: Mark-Compact 1938.6 (2098.2) -> 1923.8 (2098.0) MB, pooled: 1 MB, 31.92 / 0.04 ms (average mu = 0.305, current mu = 0.306) task; scavenge might not succeed
[20041:0x75280c000] 68684395 ms: Mark-Compact 1939.8 (2098.0) -> 1922.8 (2097.1) MB, pooled: 1 MB, 32.96 / 0.04 ms (average mu = 0.300, current mu = 0.296) allocation failure; scavenge might not succeed
FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
----- Native stack trace -----
1: 0x1043219f4 node::OOMErrorHandler(char const*, v8::OOMDetails const&) [...]
...
39: 0x188ef84e4 start [/usr/lib/dyld]
zsh: abort cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
What happened / steps to reproduce
- Run a long-lived agent conversation (this one had been active across compaction, with a large session transcript and many tool calls — screenshots, sub-agents, image reads).
- Resume that session headlessly:
cmd --resume <session-id> --yolo. - The process climbs to ~1.94 GB, GC runs become ineffective (mark-compact can't reclaim), and Node aborts with OOM.
Suspected cause
The resume path likely loads the full session transcript (which after many turns / compaction / large tool outputs — including read_file image data and agent results — grows large) into memory after the shared/global heap is already near the default Node limit. Because the process was launched by node without --max-old-space-size, it dies at the default 2 GB-ish ceiling rather than degrading gracefully. Possible contributors:
- Session transcript/context payload loaded in full on
--resume. - Tool outputs (screenshots/images/large read results from sub-agents) retained in memory instead of being streamed or evicted.
- No explicit heap sizing or soft-cap/memory-pressure handling in the resume path.
Expected behavior
A long session resume should either:
- not exceed the default heap (stream/evict large payloads), or
- surface a clear, recoverable error (e.g. "session too large, please compact/start fresh") instead of a hard
abortthat loses the work, or - raise/self-tune the heap limit (e.g.
--max-old-space-size) or chunk the resume.
Additional context
This happened on a session that had already been auto-compacted once and contained a large number of tool calls including image reads. It is the first reproducible-ish OOM in ordinary usage — no custom memory settings were in place.
Impact
- The resumed session crashes and the user loses the working context.
- Error message only points at the JS engine, not at Command Code — no hint to reduce session size or compact first.
貢獻指南
這個儲存庫沒有索引到貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 cmd --resume <session-id> --yolo 的恢復路徑開始,並使用一個包含壓縮、大型工具輸出和影像讀取的長工作階段進行重現。追蹤恢復期間工作階段 transcript 和工具結果如何載入及保留。完成的標準是工作階段避免 default heap OOM,或回報清楚且可復原的錯誤,而不是中止並遺失工作階段。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- javascript, node.js
- 領域
- cli, performance
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 活躍
- 描述清晰度
- 基本清楚
- 新手友好度
- 48/100