Resume of a large session OOMs / grinds one CPU core for ~70 min in 1.0.74 (regression vs 1.0.73; ~3–4× memory)
- 主要言語
- Shell
- スター
- 11.2k
- フォーク
- 1.9k
- 平均マージ
- 14時間 16分
- マージ済み PR(30日)
- 6
説明
### Describe the bug
After upgrading to 1.0.74, resuming a long-lived, large session (that had resumed fine daily for months) now fails. A controlled A/B on the same machine and same session, changing only the CLI version, isolates the regression to 1.0.74:
| Version | Peak RSS resuming the same ~532 MB session | Result |
|---------|--------------------------------------------|--------|
| 1.0.73 | ~2.8 GB (steady) | resumes successfully |
| 1.0.74 | ~8.5 GB (climbing) | OOMs at the default Node heap; with a raised heap it does not OOM but pegs one CPU core for ~70 minutes before becoming interactive |
1.0.74 uses ~3–4× the memory of 1.0.73 for the identical resume.
The reconstruction is also **not cached**: after the ~70-minute grind completes, nothing is persisted (session-store.db main file and WAL unchanged, session.db untouched — only a lock file and the workspace.yaml `updated_at` timestamp were written), so every subsequent resume repeats the full grind. Verified by resuming the same session twice.
Likely area: 1.0.74's subagent-timeline rework. The changelog entry "Multi-turn subagent timelines show every prompt and response in the correct order after reopening /tasks" is 1.0.74-only, and a new `_timelineReplayBuffer` structure appears in the 1.0.74 bundle (absent in 1.0.73). (The buffer itself looks bounded, so it may not be the sole cause — flagging for investigation.)
Separately, on a corporate-proxy network the remote-session export path (upload to `api.business.githubcopilot.com/agents`) returns 403 and stalls; it also reads the entire session log into memory to upload. `--no-remote-export` avoids that hang and reduces peak memory (but does not remove the ~70-min CPU grind, which is the core problem).
### Affected version
GitHub Copilot CLI 1.0.74
### Steps to reproduce the behavior
1. Have (or grow) a session with a very large events.jsonl (hundreds of MB; this one is ~532 MB with 184 checkpoints, created ~4 months ago and resumed daily).
2. Upgrade to 1.0.74.
3. `copilot --resume` and select that session (or `copilot --resume=`).
4. Observe: OOM (`FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory`), or — if the Node heap is raised — a single core pinned at ~100% for ~70 minutes before the prompt appears.
To confirm it is 1.0.74-specific, resume the same session with `--prefer-version 1.0.73`: it resumes in a fraction of the time and memory.
### Expected behavior
Resuming a large session should complete quickly and with bounded memory, as it did in 1.0.73. The resume replay / timeline reconstruction should stream rather than read+duplicate the whole events.jsonl in memory, and ideally the reconstructed state should be cached so resume is incremental rather than a full re-replay every time.
### Additional context
Evidence gathered:
- Crash reports: `trigger: OOMError`, heap total ~2716 MB (the default auto-heap on a 12 GB host).
- ~16× expansion: a 532 MB log parses to ~8.5 GB RSS under 1.0.74.
- The packaged single-executable ignores `NODE_OPTIONS`, so users cannot raise `--max-old-space-size` to work around the OOM (verified: `NODE_OPTIONS=--max-old-space-size=10` does not even fail `--version`). A supported heap override, or a percentage-of-RAM default, would help.
- Under 1.0.73 the same resume holds ~2.8 GB and succeeds.
Possibly related: #4138 (resume triggers background compaction that fails silently), #3900 (secret filtering can block the CLI UI thread — the export/replay path secret-filters the whole session), #1457 (generic JavaScript heap out of memory), #3856 / #4140 (slow /resume picker).
Suggested fixes:
1. Stream the resume replay / event-cache build instead of read-whole-file + duplicate.
2. Persist/cache the reconstructed session state (or a compacted checkpoint) so resume is incremental.
3. In the remote-export uploader, stream and batch rather than readFile → split → parse → stringify → filter → parse.
4. Fail fast on an unreachable/403 remote-session endpoint instead of reading the whole log first.
5. Honour an env/flag to raise the heap ceiling (the SEA currently ignores NODE_OPTIONS).
コントリビューションガイド
調査の方向性
--resume パスから始め、_timelineReplayBuffer と events.jsonl のリプレイ周辺にある 1.0.73 と 1.0.74 の実装を比較します。大規模なセッションで再現し、メモリと CPU を測定します。また、再構築後に session-store.db、session.db、workspace.yaml のいずれかが変更されるかどうかも確認します。完了の条件は、OOM や長時間にわたる単一コア処理なしで再開が完了し、負荷の高い完全なリプレイを不必要に繰り返さないことです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- javascript, node.js
- 領域
- cli, performance
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 48/100