github / github/copilot-cli

Node.js OOM crash after ~37 min — 31,965 leaked async libuv handles (SEA ignores NODE_OPTIONS)

Aperta
#4,686 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

area:sessions
Lingua principale
Shell
Stelle
11.2k
Fork
1.9k
Merge medio
14h 16m
PR unite (30g)
6

Descrizione

Bug: Node.js OOM crash after ~37 minutes — 31,965 leaked async libuv handles

Version: 1.0.82
Platform: Linux x86_64 (Amazon EC2, Node v24.20.0 embedded in SEA)
Observed: 2026-09-01, crash at ~37 min session uptime

Summary

Every session crashes with FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory after roughly 30–60 minutes of use. The crash is caused by a monotonic leak of async libuv handles (~32k by crash time), not by conversation-context growth. /compact and increasing NODE_OPTIONS (which the SEA ignores) do not prevent the crash.

Heap report metrics (from report.20260901.153402.2817455.0.001.json)

Metric Value
Session uptime at crash 2,203 s (~37 min)
old_space used / capacity 3,899 MB / 3,899 MB (100% saturated)
externalMemory 36 MB
mallocedMemory ~1 MB
memoryLimit 4,298 MB (~4.0 GB)
RSS at crash 4,634 MB
GC current mu 0.173 (thrashing — normally >0.5 is healthy)
CPU at crash ~292% (GC threads saturating 3 cores)

Growth rate: ~100 MB/min of old_space objects throughout the session.

libuv handle leak

At crash time, the report shows:

libuv handle count: 31,996
  type=async, is_active=true, is_referenced=false: 31,965
  type=signal, is_active=true:                         19
  • 31,965 active-but-unreferenced async handles — every one holds a JS closure/context in old_space. With no detail string (details: ""), these are likely AsyncResource / AsyncLocalStorage contexts that were created but never destroyed.
  • 19 active signal handles (vs. the ~17 signals that exist): duplicate SIGABRT, SIGIO registrations suggest repeated process.on('signal', ...) calls without corresponding removeListener — consistent with code that re-registers handlers per turn or per tool call.

The combination is a textbook listener/async-context leak: each model turn or tool execution allocates one or more async handles that are never cleaned up. After ~32k turns/tool calls the heap saturates.

Why NODE_OPTIONS and heap increases appear to help but don't

The copilot binary is a Node Single Executable Application (SEA) (NODE_SEA_BLOB present; 159 MB on disk). SEAs ignore NODE_OPTIONS entirely — empirically verified: NODE_OPTIONS="--this-is-not-a-real-flag" copilot --version exits 0 with no error.

The commandLine in the crash report confirms only Copilot's own flags are applied:

['copilot','--no-warnings','--report-on-fatalerror','--optimize-for-size','--expose-gc','copilot']

Increasing the (ignored) NODE_OPTIONS heap size is a placebo — the 4 GB cap is always in effect.

Reproduction pattern

The crash is reproducible by running a session of moderate length (>30 min) with regular tool calls (bash, view, grep). The leak appears to be per tool execution or per model turn rather than per session.

The --report-on-fatalerror flag already embedded in the binary's commandLine will produce a report.YYYYMMDD.*.json in the CWD on every crash — these reports contain full diagnostics.

Suggested investigation areas

  1. AsyncResource / AsyncLocalStorage contexts created per turn/tool that escape cleanup (search for new AsyncResource or AsyncLocalStorage usages in app.js)
  2. process.on(signal, ...) calls without process.removeListener — each session re-registration without cleanup would explain the 19 duplicate signal handles
  3. Any setInterval/setTimeout created per tool call that is never clearInterval-ed — these register as async handles too

Workaround (until fixed)

Run the Copilot JS dist under a real node with a larger heap:

~/.nvm/versions/node/v24.20.0/bin/node \
    --max-old-space-size=12288 \
    --expose-gc --no-warnings --optimize-for-size \
    ~/.cache/copilot/pkg/linux-x64/1.0.82/index.js \
    "$@"

This buys ~2 hours before the same leak exhausts 12 GB — same leak, bigger bucket.

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia riproducendo una sessione con chiamate normali agli strumenti e ispeziona il report generato report.YYYYMMDD.*.json, quindi leggi il codice per la gestione del contesto asincrono e dei segnali in app.js. Cerca i percorsi menzionati relativi a AsyncResource, AsyncLocalStorage, process.on, timer e cleanup; il lavoro è completato quando il conteggio degli handle non cresce più monotonamente e le sessioni prolungate evitano l’OOM segnalato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
javascript, node.js
Ambito
cli, devtools, performance
Tipo di issue
Bug
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Attiva
Chiarezza
Da chiarire
Idoneità per principianti
28/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.