CommandCodeAI / CommandCodeAI/command-code

OOM crash (heap out of memory) when resuming long session [1.39.2]

未關閉
#783 3 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

主要語言
沒有語言資料
星號
4k
分支
350
PR 合併指標
30 天內沒有已合併 PR

描述

Description

Running cmd --resume <session-id> --yolo on a long-running agent session crashes with a Node.js FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. The process aborts (zsh: abort) and the session is lost.

Environment
  • Command Code: 1.39.2
  • Node: v24.19.0 (/Users/shreyashphakadepawar/.nvm/versions/node/v24.19.0/bin/node)
  • OS: macOS 26.6.2 (Build 25G83), arm64
  • Invocation: cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
  • Default Node heap limit on this machine: 2240 MB (v8.getHeapStatistics().heap_size_limit)
Crash output
<--- Last few GCs --->
[20041:0x75280c000] 68684349 ms: Mark-Compact 1938.6 (2098.2) -> 1923.8 (2098.0) MB, pooled: 1 MB, 31.92 / 0.04 ms  (average mu = 0.305, current mu = 0.306) task; scavenge might not succeed
[20041:0x75280c000] 68684395 ms: Mark-Compact 1939.8 (2098.0) -> 1922.8 (2097.1) MB, pooled: 1 MB, 32.96 / 0.04 ms  (average mu = 0.300, current mu = 0.296) allocation failure; scavenge might not succeed

FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
----- Native stack trace -----
 1: 0x1043219f4 node::OOMErrorHandler(char const*, v8::OOMDetails const&) [...]
 ...
 39: 0x188ef84e4 start [/usr/lib/dyld]
zsh: abort      cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
What happened / steps to reproduce
  1. Run a long-lived agent conversation (this one had been active across compaction, with a large session transcript and many tool calls — screenshots, sub-agents, image reads).
  2. Resume that session headlessly: cmd --resume <session-id> --yolo.
  3. The process climbs to ~1.94 GB, GC runs become ineffective (mark-compact can't reclaim), and Node aborts with OOM.
Suspected cause

The resume path likely loads the full session transcript (which after many turns / compaction / large tool outputs — including read_file image data and agent results — grows large) into memory after the shared/global heap is already near the default Node limit. Because the process was launched by node without --max-old-space-size, it dies at the default 2 GB-ish ceiling rather than degrading gracefully. Possible contributors:

  • Session transcript/context payload loaded in full on --resume.
  • Tool outputs (screenshots/images/large read results from sub-agents) retained in memory instead of being streamed or evicted.
  • No explicit heap sizing or soft-cap/memory-pressure handling in the resume path.
Expected behavior

A long session resume should either:

  • not exceed the default heap (stream/evict large payloads), or
  • surface a clear, recoverable error (e.g. "session too large, please compact/start fresh") instead of a hard abort that loses the work, or
  • raise/self-tune the heap limit (e.g. --max-old-space-size) or chunk the resume.
Additional context

This happened on a session that had already been auto-compacted once and contained a large number of tool calls including image reads. It is the first reproducible-ish OOM in ordinary usage — no custom memory settings were in place.

Impact
  • The resumed session crashes and the user loses the working context.
  • Error message only points at the JS engine, not at Command Code — no hint to reduce session size or compact first.

貢獻指南

這個儲存庫沒有索引到貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

cmd --resume <session-id> --yolo 的恢復路徑開始,並使用一個包含壓縮、大型工具輸出和影像讀取的長工作階段進行重現。追蹤恢復期間工作階段 transcript 和工具結果如何載入及保留。完成的標準是工作階段避免 default heap OOM,或回報清楚且可復原的錯誤,而不是中止並遺失工作階段。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
javascript, node.js
領域
cli, performance
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
活躍
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。