v1.0.80: event-storage exhaustion retry storm drives long-running session into GC/compaction loop and Node OOM
Ninguém assumiu esta issue ainda.
- Linguagem predominante
- Shell
- Estrelas
- 11.2k
- Forks
- 1.9k
- Merge médio
- 14h 16min
- PRs com merge (30d)
- 6
Descrição
Describe the bug
A long-running active Copilot CLI session eventually reached remote event-storage exhaustion. After that, the exporter kept attempting 500-event flushes, while the process repeatedly reported memory pressure, forced GC/emergency compaction, and thousands of bridge acknowledgement timeouts. The CLI ultimately terminated with a Node out-of-memory error shown in the terminal.
This appears to connect two existing failure modes that are currently reported separately:
- #4467:
session_event_storage_exhausted - #4506: memory-pressure GC/compaction loop ending in OOM
- Also related to #4251 for unbounded memory growth in large sessions
The new observation is that remote storage exhaustion does not merely make session status unreliable: the repeated exporter failures/backlog correlate with sustained local memory pressure and eventual process death.
Environment
- Copilot CLI at failure: 1.0.80
- Embedded Node.js: v24.18.1
- OS: Linux x86_64, kernel 6.17.0-1022-azure
- Host memory: 125 GiB RAM, no swap
- Session created: 2026-08-10
- Process/resume started: 2026-08-18
- Failure: 2026-08-27 (about 8.8 days of this process, about 16.9 days of session lifetime)
- Workload: long-running interactive/autonomous software-engineering session with many tool calls and subagents
Session size at failure
events.jsonl: 934,639,321 bytes (~892 MiB)- Event records: 282,912
- Checkpoints: 32
Log evidence
Across the process log:
| Event | Count |
|---|---|
session_event_storage_exhausted |
18,909 |
Memory pressure detected - requesting garbage collection |
21,352 |
Memory pressure persists after GC - triggering emergency compaction |
473 |
timed out waiting for bridge event ack |
3,001 |
Representative final minutes:
[WARNING] Failed to submit events to Mission Control session …: 409
{"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
[WARNING] remote session batch flush failed {"event_count":500,...}
[WARNING] remote session exporter circuit opened after batch flush failure
[WARNING] Memory pressure detected - requesting garbage collection
[WARNING] Memory pressure persists after GC - triggering emergency compaction
[WARNING] timed out waiting for bridge event ack {"timeout_ms":30000}
[INFO] Compacted 17 messages, saved ~9267 tokens
[WARNING] Memory pressure detected - requesting garbage collection
...
[INFO] Compacted 38 messages, saved ~23339 tokens
[WARNING] Memory pressure detected - requesting garbage collection
The application log ends in repeated memory-pressure warnings. The fatal Node OOM text was printed by the terminating runtime to the terminal and was not persisted into the Copilot process log. No kernel OOM-killer event was recorded, so this was a Node/V8 process failure rather than the Linux kernel killing the process.
The host currently has ample free memory after restart; I did not capture /proc/<pid>/status immediately before the crash, so I cannot claim whether the final allocation failure was at the V8 heap limit or in native/external memory.
Steps to reproduce
- Keep one Copilot CLI session active/resumable for more than a week with frequent tool calls and subagent events.
- Allow
events.jsonlto grow toward 1 GB and remote Mission Control event storage to fill. - Observe repeated 409
session_event_storage_exhaustedresponses for 500-event batches. - Continue using the session.
- Observe repeated memory-pressure GC, emergency compaction, bridge ack timeouts, and eventually a Node OOM termination.
Expected behavior
- Rotate/continue/compact remote event storage before its limit is reached.
- Once the server permanently rejects a stream as exhausted, stop retaining/retrying the rejected backlog in a way that grows local memory.
- Keep local session operation and checkpointing healthy when remote export is unavailable.
- Apply bounded backoff and expose a clear degraded-export state.
- Do not repeatedly perform lossy conversation compaction unless it measurably relieves the memory pressure.
- Persist the fatal Node diagnostic report path or equivalent crash telemetry in the process log.
Actual behavior
A permanent remote quota condition produced 18,909 failed submissions, concurrent memory-pressure/compaction loops, degraded bridge responsiveness, and eventual Node OOM. Recovery required starting a new session and manually reading the old checkpoint and session directory.
Guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Direção de pesquisa
Comece pelas falhas de flush de 500 eventos do exportador de sessões remotas, pelas respostas session_event_storage_exhausted e pelas mensagens relacionadas de pressão de memória do process-log. Rastreie como os lotes rejeitados são mantidos e tentados novamente e, em seguida, verifique se uma falha permanente de exportação mantém a operação da sessão local saudável, usa um comportamento de retry limitado e registra um estado degradado claro sem levar a OOM.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- node.js
- Domínio
- backend, cli, performance
- Tipo de issue
- Bug
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Status de atividade
- Ativa
- Clareza
- Razoavelmente clara
- Facilidade para iniciantes
- 28/100