CommandCodeAI / CommandCodeAI/command-code

OOM crash (heap out of memory) when resuming long session [1.39.2]

未关闭
#783 3 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

主要语言
没有语言数据
星标
4k
派生
350
PR 合并指标
30 天内没有已合并 PR

描述

Description

Running cmd --resume <session-id> --yolo on a long-running agent session crashes with a Node.js FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. The process aborts (zsh: abort) and the session is lost.

Environment
  • Command Code: 1.39.2
  • Node: v24.19.0 (/Users/shreyashphakadepawar/.nvm/versions/node/v24.19.0/bin/node)
  • OS: macOS 26.6.2 (Build 25G83), arm64
  • Invocation: cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
  • Default Node heap limit on this machine: 2240 MB (v8.getHeapStatistics().heap_size_limit)
Crash output
<--- Last few GCs --->
[20041:0x75280c000] 68684349 ms: Mark-Compact 1938.6 (2098.2) -> 1923.8 (2098.0) MB, pooled: 1 MB, 31.92 / 0.04 ms  (average mu = 0.305, current mu = 0.306) task; scavenge might not succeed
[20041:0x75280c000] 68684395 ms: Mark-Compact 1939.8 (2098.0) -> 1922.8 (2097.1) MB, pooled: 1 MB, 32.96 / 0.04 ms  (average mu = 0.300, current mu = 0.296) allocation failure; scavenge might not succeed

FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
----- Native stack trace -----
 1: 0x1043219f4 node::OOMErrorHandler(char const*, v8::OOMDetails const&) [...]
 ...
 39: 0x188ef84e4 start [/usr/lib/dyld]
zsh: abort      cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
What happened / steps to reproduce
  1. Run a long-lived agent conversation (this one had been active across compaction, with a large session transcript and many tool calls — screenshots, sub-agents, image reads).
  2. Resume that session headlessly: cmd --resume <session-id> --yolo.
  3. The process climbs to ~1.94 GB, GC runs become ineffective (mark-compact can't reclaim), and Node aborts with OOM.
Suspected cause

The resume path likely loads the full session transcript (which after many turns / compaction / large tool outputs — including read_file image data and agent results — grows large) into memory after the shared/global heap is already near the default Node limit. Because the process was launched by node without --max-old-space-size, it dies at the default 2 GB-ish ceiling rather than degrading gracefully. Possible contributors:

  • Session transcript/context payload loaded in full on --resume.
  • Tool outputs (screenshots/images/large read results from sub-agents) retained in memory instead of being streamed or evicted.
  • No explicit heap sizing or soft-cap/memory-pressure handling in the resume path.
Expected behavior

A long session resume should either:

  • not exceed the default heap (stream/evict large payloads), or
  • surface a clear, recoverable error (e.g. "session too large, please compact/start fresh") instead of a hard abort that loses the work, or
  • raise/self-tune the heap limit (e.g. --max-old-space-size) or chunk the resume.
Additional context

This happened on a session that had already been auto-compacted once and contained a large number of tool calls including image reads. It is the first reproducible-ish OOM in ordinary usage — no custom memory settings were in place.

Impact
  • The resumed session crashes and the user loses the working context.
  • Error message only points at the JS engine, not at Command Code — no hint to reduce session size or compact first.

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

cmd --resume <session-id> --yolo 恢复路径开始,并使用一个包含压缩、大型工具输出和图像读取的长会话进行复现。跟踪恢复期间会话 transcript 和工具结果是如何加载和保留的。完成的标准是会话避免 default heap OOM,或者报告清晰且可恢复的错误,而不是中止并丢失会话。

由索引模型根据 Issue 内容生成。

评估

技术栈
javascript, node.js
领域
cli, performance
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
活跃
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。