github / github/copilot-cli

Resume of a large session OOMs / grinds one CPU core for ~70 min in 1.0.74 (regression vs 1.0.73; ~3–4× memory)

Open
#4,251 4 comments 1 reaction 0 assignees View on GitHub
area:sessions
Dominant language
Shell
Stars
11.2k
Forks
1.9k
Avg merge
14h 16m
Merged PRs (30d)
6

Description

### Describe the bug

After upgrading to 1.0.74, resuming a long-lived, large session (that had resumed fine daily for months) now fails. A controlled A/B on the same machine and same session, changing only the CLI version, isolates the regression to 1.0.74:

| Version | Peak RSS resuming the same ~532 MB session | Result |
|---------|--------------------------------------------|--------|
| 1.0.73 | ~2.8 GB (steady) | resumes successfully |
| 1.0.74 | ~8.5 GB (climbing) | OOMs at the default Node heap; with a raised heap it does not OOM but pegs one CPU core for ~70 minutes before becoming interactive |

1.0.74 uses ~3–4× the memory of 1.0.73 for the identical resume.

The reconstruction is also **not cached**: after the ~70-minute grind completes, nothing is persisted (session-store.db main file and WAL unchanged, session.db untouched — only a lock file and the workspace.yaml `updated_at` timestamp were written), so every subsequent resume repeats the full grind. Verified by resuming the same session twice.

Likely area: 1.0.74's subagent-timeline rework. The changelog entry "Multi-turn subagent timelines show every prompt and response in the correct order after reopening /tasks" is 1.0.74-only, and a new `_timelineReplayBuffer` structure appears in the 1.0.74 bundle (absent in 1.0.73). (The buffer itself looks bounded, so it may not be the sole cause — flagging for investigation.)

Separately, on a corporate-proxy network the remote-session export path (upload to `api.business.githubcopilot.com/agents`) returns 403 and stalls; it also reads the entire session log into memory to upload. `--no-remote-export` avoids that hang and reduces peak memory (but does not remove the ~70-min CPU grind, which is the core problem).

### Affected version

GitHub Copilot CLI 1.0.74

### Steps to reproduce the behavior

1. Have (or grow) a session with a very large events.jsonl (hundreds of MB; this one is ~532 MB with 184 checkpoints, created ~4 months ago and resumed daily).
2. Upgrade to 1.0.74.
3. `copilot --resume` and select that session (or `copilot --resume=`).
4. Observe: OOM (`FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory`), or — if the Node heap is raised — a single core pinned at ~100% for ~70 minutes before the prompt appears.

To confirm it is 1.0.74-specific, resume the same session with `--prefer-version 1.0.73`: it resumes in a fraction of the time and memory.

### Expected behavior

Resuming a large session should complete quickly and with bounded memory, as it did in 1.0.73. The resume replay / timeline reconstruction should stream rather than read+duplicate the whole events.jsonl in memory, and ideally the reconstructed state should be cached so resume is incremental rather than a full re-replay every time.

### Additional context

Evidence gathered:
- Crash reports: `trigger: OOMError`, heap total ~2716 MB (the default auto-heap on a 12 GB host).
- ~16× expansion: a 532 MB log parses to ~8.5 GB RSS under 1.0.74.
- The packaged single-executable ignores `NODE_OPTIONS`, so users cannot raise `--max-old-space-size` to work around the OOM (verified: `NODE_OPTIONS=--max-old-space-size=10` does not even fail `--version`). A supported heap override, or a percentage-of-RAM default, would help.
- Under 1.0.73 the same resume holds ~2.8 GB and succeeds.

Possibly related: #4138 (resume triggers background compaction that fails silently), #3900 (secret filtering can block the CLI UI thread — the export/replay path secret-filters the whole session), #1457 (generic JavaScript heap out of memory), #3856 / #4140 (slow /resume picker).

Suggested fixes:
1. Stream the resume replay / event-cache build instead of read-whole-file + duplicate.
2. Persist/cache the reconstructed session state (or a compacted checkpoint) so resume is incremental.
3. In the remote-export uploader, stream and batch rather than readFile → split → parse → stringify → filter → parse.
4. Fail fast on an unreachable/403 remote-session endpoint instead of reading the whole log first.
5. Honour an env/flag to raise the heap ceiling (the SEA currently ignores NODE_OPTIONS).

Contributor guide

Open the contributing guide

Research direction

Start with the --resume path and compare the 1.0.73 and 1.0.74 implementations around the _timelineReplayBuffer and events.jsonl replay. Reproduce with the large session, measuring memory and CPU; also check whether session-store.db, session.db, or workspace.yaml change after reconstruction. Done means resuming completes without OOM or prolonged single-core processing and does not repeat the full expensive replay unnecessarily.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript, node.js
Domain
cli, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.