v1.0.80: event-storage exhaustion retry storm drives long-running session into GC/compaction loop and Node OOM
Nessuno ha ancora preso questa issue.
- Lingua principale
- Shell
- Stelle
- 11.2k
- Fork
- 1.9k
- Merge medio
- 14h 16m
- PR unite (30g)
- 6
Descrizione
Describe the bug
A long-running active Copilot CLI session eventually reached remote event-storage exhaustion. After that, the exporter kept attempting 500-event flushes, while the process repeatedly reported memory pressure, forced GC/emergency compaction, and thousands of bridge acknowledgement timeouts. The CLI ultimately terminated with a Node out-of-memory error shown in the terminal.
This appears to connect two existing failure modes that are currently reported separately:
- #4467:
session_event_storage_exhausted - #4506: memory-pressure GC/compaction loop ending in OOM
- Also related to #4251 for unbounded memory growth in large sessions
The new observation is that remote storage exhaustion does not merely make session status unreliable: the repeated exporter failures/backlog correlate with sustained local memory pressure and eventual process death.
Environment
- Copilot CLI at failure: 1.0.80
- Embedded Node.js: v24.18.1
- OS: Linux x86_64, kernel 6.17.0-1022-azure
- Host memory: 125 GiB RAM, no swap
- Session created: 2026-08-10
- Process/resume started: 2026-08-18
- Failure: 2026-08-27 (about 8.8 days of this process, about 16.9 days of session lifetime)
- Workload: long-running interactive/autonomous software-engineering session with many tool calls and subagents
Session size at failure
events.jsonl: 934,639,321 bytes (~892 MiB)- Event records: 282,912
- Checkpoints: 32
Log evidence
Across the process log:
| Event | Count |
|---|---|
session_event_storage_exhausted |
18,909 |
Memory pressure detected - requesting garbage collection |
21,352 |
Memory pressure persists after GC - triggering emergency compaction |
473 |
timed out waiting for bridge event ack |
3,001 |
Representative final minutes:
[WARNING] Failed to submit events to Mission Control session …: 409
{"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
[WARNING] remote session batch flush failed {"event_count":500,...}
[WARNING] remote session exporter circuit opened after batch flush failure
[WARNING] Memory pressure detected - requesting garbage collection
[WARNING] Memory pressure persists after GC - triggering emergency compaction
[WARNING] timed out waiting for bridge event ack {"timeout_ms":30000}
[INFO] Compacted 17 messages, saved ~9267 tokens
[WARNING] Memory pressure detected - requesting garbage collection
...
[INFO] Compacted 38 messages, saved ~23339 tokens
[WARNING] Memory pressure detected - requesting garbage collection
The application log ends in repeated memory-pressure warnings. The fatal Node OOM text was printed by the terminating runtime to the terminal and was not persisted into the Copilot process log. No kernel OOM-killer event was recorded, so this was a Node/V8 process failure rather than the Linux kernel killing the process.
The host currently has ample free memory after restart; I did not capture /proc/<pid>/status immediately before the crash, so I cannot claim whether the final allocation failure was at the V8 heap limit or in native/external memory.
Steps to reproduce
- Keep one Copilot CLI session active/resumable for more than a week with frequent tool calls and subagent events.
- Allow
events.jsonlto grow toward 1 GB and remote Mission Control event storage to fill. - Observe repeated 409
session_event_storage_exhaustedresponses for 500-event batches. - Continue using the session.
- Observe repeated memory-pressure GC, emergency compaction, bridge ack timeouts, and eventually a Node OOM termination.
Expected behavior
- Rotate/continue/compact remote event storage before its limit is reached.
- Once the server permanently rejects a stream as exhausted, stop retaining/retrying the rejected backlog in a way that grows local memory.
- Keep local session operation and checkpointing healthy when remote export is unavailable.
- Apply bounded backoff and expose a clear degraded-export state.
- Do not repeatedly perform lossy conversation compaction unless it measurably relieves the memory pressure.
- Persist the fatal Node diagnostic report path or equivalent crash telemetry in the process log.
Actual behavior
A permanent remote quota condition produced 18,909 failed submissions, concurrent memory-pressure/compaction loops, degraded bridge responsiveness, and eventual Node OOM. Recovery required starting a new session and manually reading the old checkpoint and session directory.
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia dai fallimenti di flush da 500 eventi dell’esportatore di sessioni remote, dalle risposte session_event_storage_exhausted e dai messaggi correlati di pressione della memoria di process-log. Traccia il modo in cui i batch rifiutati vengono conservati e ritentati, quindi verifica che un errore permanente di esportazione lasci sana l’operazione della sessione locale, utilizzi un comportamento di retry limitato e registri uno stato degradato chiaro senza portare a OOM.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- node.js
- Ambito
- backend, cli, performance
- Tipo di issue
- Bug
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 28/100