github / github/copilot-cli

Node.js OOM crash after ~37 min — 31,965 leaked async libuv handles (SEA ignores NODE_OPTIONS)

Open
#4,686 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

area:sessions
Dominant language
Shell
Stars
11.2k
Forks
1.9k
Avg merge
14h 16m
Merged PRs (30d)
6

Description

Bug: Node.js OOM crash after ~37 minutes — 31,965 leaked async libuv handles

Version: 1.0.82
Platform: Linux x86_64 (Amazon EC2, Node v24.20.0 embedded in SEA)
Observed: 2026-09-01, crash at ~37 min session uptime

Summary

Every session crashes with FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory after roughly 30–60 minutes of use. The crash is caused by a monotonic leak of async libuv handles (~32k by crash time), not by conversation-context growth. /compact and increasing NODE_OPTIONS (which the SEA ignores) do not prevent the crash.

Heap report metrics (from report.20260901.153402.2817455.0.001.json)

Metric Value
Session uptime at crash 2,203 s (~37 min)
old_space used / capacity 3,899 MB / 3,899 MB (100% saturated)
externalMemory 36 MB
mallocedMemory ~1 MB
memoryLimit 4,298 MB (~4.0 GB)
RSS at crash 4,634 MB
GC current mu 0.173 (thrashing — normally >0.5 is healthy)
CPU at crash ~292% (GC threads saturating 3 cores)

Growth rate: ~100 MB/min of old_space objects throughout the session.

libuv handle leak

At crash time, the report shows:

libuv handle count: 31,996
  type=async, is_active=true, is_referenced=false: 31,965
  type=signal, is_active=true:                         19
  • 31,965 active-but-unreferenced async handles — every one holds a JS closure/context in old_space. With no detail string (details: ""), these are likely AsyncResource / AsyncLocalStorage contexts that were created but never destroyed.
  • 19 active signal handles (vs. the ~17 signals that exist): duplicate SIGABRT, SIGIO registrations suggest repeated process.on('signal', ...) calls without corresponding removeListener — consistent with code that re-registers handlers per turn or per tool call.

The combination is a textbook listener/async-context leak: each model turn or tool execution allocates one or more async handles that are never cleaned up. After ~32k turns/tool calls the heap saturates.

Why NODE_OPTIONS and heap increases appear to help but don't

The copilot binary is a Node Single Executable Application (SEA) (NODE_SEA_BLOB present; 159 MB on disk). SEAs ignore NODE_OPTIONS entirely — empirically verified: NODE_OPTIONS="--this-is-not-a-real-flag" copilot --version exits 0 with no error.

The commandLine in the crash report confirms only Copilot's own flags are applied:

['copilot','--no-warnings','--report-on-fatalerror','--optimize-for-size','--expose-gc','copilot']

Increasing the (ignored) NODE_OPTIONS heap size is a placebo — the 4 GB cap is always in effect.

Reproduction pattern

The crash is reproducible by running a session of moderate length (>30 min) with regular tool calls (bash, view, grep). The leak appears to be per tool execution or per model turn rather than per session.

The --report-on-fatalerror flag already embedded in the binary's commandLine will produce a report.YYYYMMDD.*.json in the CWD on every crash — these reports contain full diagnostics.

Suggested investigation areas

  1. AsyncResource / AsyncLocalStorage contexts created per turn/tool that escape cleanup (search for new AsyncResource or AsyncLocalStorage usages in app.js)
  2. process.on(signal, ...) calls without process.removeListener — each session re-registration without cleanup would explain the 19 duplicate signal handles
  3. Any setInterval/setTimeout created per tool call that is never clearInterval-ed — these register as async handles too

Workaround (until fixed)

Run the Copilot JS dist under a real node with a larger heap:

~/.nvm/versions/node/v24.20.0/bin/node \
    --max-old-space-size=12288 \
    --expose-gc --no-warnings --optimize-for-size \
    ~/.cache/copilot/pkg/linux-x64/1.0.82/index.js \
    "$@"

This buys ~2 hours before the same leak exhausts 12 GB — same leak, bigger bucket.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing a session with regular tool calls and inspect the generated report.YYYYMMDD.*.json, then read the async-context and signal-handling code in app.js. Search the mentioned AsyncResource, AsyncLocalStorage, process.on, timer, and cleanup paths; done means the handle count no longer grows monotonically and extended sessions avoid the reported OOM.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript, node.js
Domain
cli, devtools, performance
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.