Resume of a large session OOMs / grinds one CPU core for ~70 min in 1.0.74 (regression vs 1.0.73; ~3–4× memory)
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Mô tả
### Describe the bug
After upgrading to 1.0.74, resuming a long-lived, large session (that had resumed fine daily for months) now fails. A controlled A/B on the same machine and same session, changing only the CLI version, isolates the regression to 1.0.74:
| Version | Peak RSS resuming the same ~532 MB session | Result |
|---------|--------------------------------------------|--------|
| 1.0.73 | ~2.8 GB (steady) | resumes successfully |
| 1.0.74 | ~8.5 GB (climbing) | OOMs at the default Node heap; with a raised heap it does not OOM but pegs one CPU core for ~70 minutes before becoming interactive |
1.0.74 uses ~3–4× the memory of 1.0.73 for the identical resume.
The reconstruction is also **not cached**: after the ~70-minute grind completes, nothing is persisted (session-store.db main file and WAL unchanged, session.db untouched — only a lock file and the workspace.yaml `updated_at` timestamp were written), so every subsequent resume repeats the full grind. Verified by resuming the same session twice.
Likely area: 1.0.74's subagent-timeline rework. The changelog entry "Multi-turn subagent timelines show every prompt and response in the correct order after reopening /tasks" is 1.0.74-only, and a new `_timelineReplayBuffer` structure appears in the 1.0.74 bundle (absent in 1.0.73). (The buffer itself looks bounded, so it may not be the sole cause — flagging for investigation.)
Separately, on a corporate-proxy network the remote-session export path (upload to `api.business.githubcopilot.com/agents`) returns 403 and stalls; it also reads the entire session log into memory to upload. `--no-remote-export` avoids that hang and reduces peak memory (but does not remove the ~70-min CPU grind, which is the core problem).
### Affected version
GitHub Copilot CLI 1.0.74
### Steps to reproduce the behavior
1. Have (or grow) a session with a very large events.jsonl (hundreds of MB; this one is ~532 MB with 184 checkpoints, created ~4 months ago and resumed daily).
2. Upgrade to 1.0.74.
3. `copilot --resume` and select that session (or `copilot --resume=`).
4. Observe: OOM (`FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory`), or — if the Node heap is raised — a single core pinned at ~100% for ~70 minutes before the prompt appears.
To confirm it is 1.0.74-specific, resume the same session with `--prefer-version 1.0.73`: it resumes in a fraction of the time and memory.
### Expected behavior
Resuming a large session should complete quickly and with bounded memory, as it did in 1.0.73. The resume replay / timeline reconstruction should stream rather than read+duplicate the whole events.jsonl in memory, and ideally the reconstructed state should be cached so resume is incremental rather than a full re-replay every time.
### Additional context
Evidence gathered:
- Crash reports: `trigger: OOMError`, heap total ~2716 MB (the default auto-heap on a 12 GB host).
- ~16× expansion: a 532 MB log parses to ~8.5 GB RSS under 1.0.74.
- The packaged single-executable ignores `NODE_OPTIONS`, so users cannot raise `--max-old-space-size` to work around the OOM (verified: `NODE_OPTIONS=--max-old-space-size=10` does not even fail `--version`). A supported heap override, or a percentage-of-RAM default, would help.
- Under 1.0.73 the same resume holds ~2.8 GB and succeeds.
Possibly related: #4138 (resume triggers background compaction that fails silently), #3900 (secret filtering can block the CLI UI thread — the export/replay path secret-filters the whole session), #1457 (generic JavaScript heap out of memory), #3856 / #4140 (slow /resume picker).
Suggested fixes:
1. Stream the resume replay / event-cache build instead of read-whole-file + duplicate.
2. Persist/cache the reconstructed session state (or a compacted checkpoint) so resume is incremental.
3. In the remote-export uploader, stream and batch rather than readFile → split → parse → stringify → filter → parse.
4. Fail fast on an unreachable/403 remote-session endpoint instead of reading the whole log first.
5. Honour an env/flag to raise the heap ceiling (the SEA currently ignores NODE_OPTIONS).
Hướng dẫn đóng góp
Hướng nghiên cứu
Bắt đầu với đường dẫn --resume và so sánh các phần triển khai của 1.0.73 và 1.0.74 xung quanh _timelineReplayBuffer và việc replay events.jsonl. Tái hiện với session lớn, đo bộ nhớ và CPU; đồng thời kiểm tra xem session-store.db, session.db hoặc workspace.yaml có thay đổi sau khi tái dựng hay không. Được xem là hoàn tất khi việc tiếp tục chạy hoàn thành mà không gặp OOM hoặc xử lý kéo dài trên một lõi đơn, đồng thời không lặp lại toàn bộ replay tốn kém một cách không cần thiết.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- javascript, node.js
- Lĩnh vực
- cli, performance
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 48/100