github / github/copilot-cli

v1.0.80: event-storage exhaustion retry storm drives long-running session into GC/compaction loop and Node OOM

Aberta
#4,639 2 comentários 2 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

triage
Linguagem predominante
Shell
Estrelas
11.2k
Forks
1.9k
Merge médio
14h 16min
PRs com merge (30d)
6

Descrição

Describe the bug

A long-running active Copilot CLI session eventually reached remote event-storage exhaustion. After that, the exporter kept attempting 500-event flushes, while the process repeatedly reported memory pressure, forced GC/emergency compaction, and thousands of bridge acknowledgement timeouts. The CLI ultimately terminated with a Node out-of-memory error shown in the terminal.

This appears to connect two existing failure modes that are currently reported separately:

  • #4467: session_event_storage_exhausted
  • #4506: memory-pressure GC/compaction loop ending in OOM
  • Also related to #4251 for unbounded memory growth in large sessions

The new observation is that remote storage exhaustion does not merely make session status unreliable: the repeated exporter failures/backlog correlate with sustained local memory pressure and eventual process death.

Environment
  • Copilot CLI at failure: 1.0.80
  • Embedded Node.js: v24.18.1
  • OS: Linux x86_64, kernel 6.17.0-1022-azure
  • Host memory: 125 GiB RAM, no swap
  • Session created: 2026-08-10
  • Process/resume started: 2026-08-18
  • Failure: 2026-08-27 (about 8.8 days of this process, about 16.9 days of session lifetime)
  • Workload: long-running interactive/autonomous software-engineering session with many tool calls and subagents
Session size at failure
  • events.jsonl: 934,639,321 bytes (~892 MiB)
  • Event records: 282,912
  • Checkpoints: 32
Log evidence

Across the process log:

Event Count
session_event_storage_exhausted 18,909
Memory pressure detected - requesting garbage collection 21,352
Memory pressure persists after GC - triggering emergency compaction 473
timed out waiting for bridge event ack 3,001

Representative final minutes:

[WARNING] Failed to submit events to Mission Control session …: 409
{"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
[WARNING] remote session batch flush failed {"event_count":500,...}
[WARNING] remote session exporter circuit opened after batch flush failure
[WARNING] Memory pressure detected - requesting garbage collection
[WARNING] Memory pressure persists after GC - triggering emergency compaction
[WARNING] timed out waiting for bridge event ack {"timeout_ms":30000}
[INFO] Compacted 17 messages, saved ~9267 tokens
[WARNING] Memory pressure detected - requesting garbage collection
...
[INFO] Compacted 38 messages, saved ~23339 tokens
[WARNING] Memory pressure detected - requesting garbage collection

The application log ends in repeated memory-pressure warnings. The fatal Node OOM text was printed by the terminating runtime to the terminal and was not persisted into the Copilot process log. No kernel OOM-killer event was recorded, so this was a Node/V8 process failure rather than the Linux kernel killing the process.

The host currently has ample free memory after restart; I did not capture /proc/<pid>/status immediately before the crash, so I cannot claim whether the final allocation failure was at the V8 heap limit or in native/external memory.

Steps to reproduce
  1. Keep one Copilot CLI session active/resumable for more than a week with frequent tool calls and subagent events.
  2. Allow events.jsonl to grow toward 1 GB and remote Mission Control event storage to fill.
  3. Observe repeated 409 session_event_storage_exhausted responses for 500-event batches.
  4. Continue using the session.
  5. Observe repeated memory-pressure GC, emergency compaction, bridge ack timeouts, and eventually a Node OOM termination.
Expected behavior
  • Rotate/continue/compact remote event storage before its limit is reached.
  • Once the server permanently rejects a stream as exhausted, stop retaining/retrying the rejected backlog in a way that grows local memory.
  • Keep local session operation and checkpointing healthy when remote export is unavailable.
  • Apply bounded backoff and expose a clear degraded-export state.
  • Do not repeatedly perform lossy conversation compaction unless it measurably relieves the memory pressure.
  • Persist the fatal Node diagnostic report path or equivalent crash telemetry in the process log.
Actual behavior

A permanent remote quota condition produced 18,909 failed submissions, concurrent memory-pressure/compaction loops, degraded bridge responsiveness, and eventual Node OOM. Recovery required starting a new session and manually reading the old checkpoint and session directory.

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece pelas falhas de flush de 500 eventos do exportador de sessões remotas, pelas respostas session_event_storage_exhausted e pelas mensagens relacionadas de pressão de memória do process-log. Rastreie como os lotes rejeitados são mantidos e tentados novamente e, em seguida, verifique se uma falha permanente de exportação mantém a operação da sessão local saudável, usa um comportamento de retry limitado e registra um estado degradado claro sem levar a OOM.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
node.js
Domínio
backend, cli, performance
Tipo de issue
Bug
Dificuldade
5/5
Tempo estimado
Mais de uma semana
Status de atividade
Ativa
Clareza
Razoavelmente clara
Facilidade para iniciantes
28/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.