github / github/copilot-cli

Resume of a large session OOMs / grinds one CPU core for ~70 min in 1.0.74 (regression vs 1.0.73; ~3–4× memory)

Đang mở
#4,251 4 bình luận 1 reaction 0 người được giao Xem trên GitHub
area:sessions
Ngôn ngữ chính
Shell
Star
11.2k
Fork
1.9k
Merge trung bình
14 giờ 16 phút
Pull request đã merge (30 ngày)
6

Mô tả

### Describe the bug

After upgrading to 1.0.74, resuming a long-lived, large session (that had resumed fine daily for months) now fails. A controlled A/B on the same machine and same session, changing only the CLI version, isolates the regression to 1.0.74:

| Version | Peak RSS resuming the same ~532 MB session | Result |
|---------|--------------------------------------------|--------|
| 1.0.73 | ~2.8 GB (steady) | resumes successfully |
| 1.0.74 | ~8.5 GB (climbing) | OOMs at the default Node heap; with a raised heap it does not OOM but pegs one CPU core for ~70 minutes before becoming interactive |

1.0.74 uses ~3–4× the memory of 1.0.73 for the identical resume.

The reconstruction is also **not cached**: after the ~70-minute grind completes, nothing is persisted (session-store.db main file and WAL unchanged, session.db untouched — only a lock file and the workspace.yaml `updated_at` timestamp were written), so every subsequent resume repeats the full grind. Verified by resuming the same session twice.

Likely area: 1.0.74's subagent-timeline rework. The changelog entry "Multi-turn subagent timelines show every prompt and response in the correct order after reopening /tasks" is 1.0.74-only, and a new `_timelineReplayBuffer` structure appears in the 1.0.74 bundle (absent in 1.0.73). (The buffer itself looks bounded, so it may not be the sole cause — flagging for investigation.)

Separately, on a corporate-proxy network the remote-session export path (upload to `api.business.githubcopilot.com/agents`) returns 403 and stalls; it also reads the entire session log into memory to upload. `--no-remote-export` avoids that hang and reduces peak memory (but does not remove the ~70-min CPU grind, which is the core problem).

### Affected version

GitHub Copilot CLI 1.0.74

### Steps to reproduce the behavior

1. Have (or grow) a session with a very large events.jsonl (hundreds of MB; this one is ~532 MB with 184 checkpoints, created ~4 months ago and resumed daily).
2. Upgrade to 1.0.74.
3. `copilot --resume` and select that session (or `copilot --resume=`).
4. Observe: OOM (`FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory`), or — if the Node heap is raised — a single core pinned at ~100% for ~70 minutes before the prompt appears.

To confirm it is 1.0.74-specific, resume the same session with `--prefer-version 1.0.73`: it resumes in a fraction of the time and memory.

### Expected behavior

Resuming a large session should complete quickly and with bounded memory, as it did in 1.0.73. The resume replay / timeline reconstruction should stream rather than read+duplicate the whole events.jsonl in memory, and ideally the reconstructed state should be cached so resume is incremental rather than a full re-replay every time.

### Additional context

Evidence gathered:
- Crash reports: `trigger: OOMError`, heap total ~2716 MB (the default auto-heap on a 12 GB host).
- ~16× expansion: a 532 MB log parses to ~8.5 GB RSS under 1.0.74.
- The packaged single-executable ignores `NODE_OPTIONS`, so users cannot raise `--max-old-space-size` to work around the OOM (verified: `NODE_OPTIONS=--max-old-space-size=10` does not even fail `--version`). A supported heap override, or a percentage-of-RAM default, would help.
- Under 1.0.73 the same resume holds ~2.8 GB and succeeds.

Possibly related: #4138 (resume triggers background compaction that fails silently), #3900 (secret filtering can block the CLI UI thread — the export/replay path secret-filters the whole session), #1457 (generic JavaScript heap out of memory), #3856 / #4140 (slow /resume picker).

Suggested fixes:
1. Stream the resume replay / event-cache build instead of read-whole-file + duplicate.
2. Persist/cache the reconstructed session state (or a compacted checkpoint) so resume is incremental.
3. In the remote-export uploader, stream and batch rather than readFile → split → parse → stringify → filter → parse.
4. Fail fast on an unreachable/403 remote-session endpoint instead of reading the whole log first.
5. Honour an env/flag to raise the heap ceiling (the SEA currently ignores NODE_OPTIONS).

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu với đường dẫn --resume và so sánh các phần triển khai của 1.0.73 và 1.0.74 xung quanh _timelineReplayBuffer và việc replay events.jsonl. Tái hiện với session lớn, đo bộ nhớ và CPU; đồng thời kiểm tra xem session-store.db, session.db hoặc workspace.yaml có thay đổi sau khi tái dựng hay không. Được xem là hoàn tất khi việc tiếp tục chạy hoàn thành mà không gặp OOM hoặc xử lý kéo dài trên một lõi đơn, đồng thời không lặp lại toàn bộ replay tốn kém một cách không cần thiết.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
javascript, node.js
Lĩnh vực
cli, performance
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
48/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.