github / github/copilot-cli

Memory-pressure watchdog force-compacts the conversation at 23% context usage, recovers 0.003% of tokens, then loops until OOM

オープン
#4,506 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

area:context-memory area:sessions
主要言語
Shell
スター
11.2k
フォーク
1.9k
平均マージ
14時間 16分
マージ済み PR(30日)
6

説明

Summary

A long-running session compacted repeatedly and aggressively while context usage was low (~23% of a 400k window). The trigger was not context pressure — it was the process memory-pressure watchdog, which force-compacts the conversation whenever process memory is high, without checking whether the conversation is actually the memory being consumed.

In this session the memory was dominated by the session's own events.jsonl (381 MB on disk, held parsed in memory), which compaction does not meaningfully shrink. So compaction ran, freed almost nothing, and re-triggered on the next tick — indefinitely. The process eventually OOM'd and died.

Related but distinct from #2132 (which is about parallel background agents). This report is about the watchdog's trigger logic and the unbounded growth of events.jsonl.

Version / environment
  • Copilot CLI 1.0.80 (@github/copilot-linux-x64)
  • Linux x64, 15 GB RAM / 48 GB swap
  • claude-opus-5, effortLevel: max, contextTier: long_context
  • max_context_window_tokens: 400000
  • V8 heap limit: 2144 MB
The core bug: compaction used to relieve memory it cannot relieve

Log lines from the affected session:

[WARNING] Memory pressure detected - requesting garbage collection
[DEBUG]   Evicted 3 transient events after compaction (93493 → 93490)
[DEBUG]   Memory pressure persists but compaction already in progress
[DEBUG]   Memory pressure persists but conversation is too small for compaction to help (~38115 < 131072 estimated bytes)

The second line is the clearest evidence:

Compaction ran at 93,493 tokens against a 400,000-token window (23% used) and recovered 3 tokens (0.003%).

It was invoked for memory reasons, did nothing for memory, and destroyed conversation history as a side effect.

Frequency within a single 200 MB slice of one session log:

line count
Memory pressure detected - requesting garbage collection 54
Memory pressure persists but compaction already in progress 42
Memory pressure persists but conversation is too small... 4+

The only guard is a 128 KB floor:

conversation is too small for compaction to help (~38115 < 131072 estimated bytes)

Above that floor the CLI compacts blindly, with no check that the conversation is a meaningful share of process memory, and no cooldown when a compaction proves ineffective. This session sat just above the floor for hours and compacted on essentially every turn.

Process state before the crash

Single CLI process, uptime 1d 13h:

metric value
RSS 1.3 GB
Swap 10.5 GB
Committed ~11.8 GB
V8 heap limit 2144 MB
Conversation ~38 KB
events.jsonl 381,627,178 bytes

~11.8 GB committed against a 38 KB conversation. Committed memory far exceeds the V8 heap limit, confirming much of it is outside the JS heap and unreclaimable by compaction under any circumstances. The process later died with:

Writing Node.js report to file: report.20260816.130754.7404.0.001.json
1: 0x74eae8 node::OOMErrorHandler(char const*, v8::OOMDetails const&)
Contributing factor: logLevel: "all" emits unbounded HTML tokenizer traces

With "logLevel": "all" (a user setting, not the default), the native runtime forwards html5ever internal tracing into the application log — one line per HTML token during web_fetch page parsing:

[DEBUG] [rust:log] processing CharacterTokens(NotSplit, Tendril<UTF8>...
[DEBUG] [rust:log] processing TagToken(Tag { kind: StartTag, name: A...
[DEBUG] [rust:log] char ref tokenizer stepping in state Named {
[DEBUG] [rust:log] char ref tokenizer stepping in state Octothorpe {

In a 200 MB sample, 62,771 of 714,788 lines were rust:log, dominated by tokenizer steps.

log size DEBUG lines avg bytes/line
session A 2,538,823,191 (2.4 GB) 4,359,820 ~580
session B 1,874,035,801 (1.8 GB) 521,347 ~3,600
session C 1,186,764,551 (1.1 GB) 174,819 ~6,800

~/.copilot/logs reached 5.8 GB. Sessions B and C average 3.6–6.8 KB per line, so large payloads are also being dumped verbatim at debug level. A logging setting should not be able to consume 5.8 GB of disk or contribute to an OOM.

For comparison, on the same machine:

--log-level output for an identical trivial run
default 1,281 bytes, 9 INFO, 0 DEBUG
all 2.4 GB, 4,359,820 DEBUG lines
Reproduction
  1. Set "logLevel": "all" in ~/.copilot/settings.json.
  2. Run a session for many hours, including several web_fetch calls against large HTML pages.
  3. Watch ~/.copilot/session-state/<id>/events.jsonl grow into the hundreds of MB.
  4. Observe process committed memory climb to several GB while the conversation stays small.
  5. Observe repeated compactions in the UI at low context usage, with the Memory pressure lines above in the log, ending in an OOM.
Expected

Compaction is driven by context-window pressure. If memory pressure is a separate trigger, it should fire only when the conversation is actually the memory being consumed, should not repeat when it has proven ineffective, and should be visible to the user.

Actual

Compaction fires at 23% context usage, recovers 0.003% of tokens, repeats indefinitely because the memory it targets is held elsewhere, and the process eventually OOMs.

Suggested fixes
  1. Before compacting for memory reasons, verify the conversation is a material fraction of process memory; if not, skip compaction and surface the real cause.
  2. Add a cooldown and an effectiveness check — if a memory-triggered compaction does not reduce measured pressure, do not retry every tick.
  3. Measure JS-heap usage attributable to the conversation separately from total RSS, so the trigger reflects something compaction can influence.
  4. Bound events.jsonl: rotate, or stream it rather than holding the parsed history resident.
  5. Scope the tracing filter to copilot_* targets so third-party crates (html5ever, hyper_util) never reach the user log, even at all. Add per-line payload caps and log rotation.
  6. Distinguish memory-triggered compaction from context-triggered compaction in the UI. A silent context wipe with no stated reason is indistinguishable from a bug.
Why memory-triggered compaction is the wrong design, independent of this bug

Even with the loop fixed, using compaction as a memory-relief mechanism is questionable. Some reasons:

It conflates two unrelated resources. The context window is a protocol limit — how many tokens the model will accept. Process memory is a runtime limit — how many bytes the host can hold. Compaction exists to manage the first. Applying it to the second assumes the two are proportional. Nothing enforces that, and this report is a case where they diverged by five orders of magnitude: a 38 KB conversation inside an 11.8 GB process.

It spends the user's data to solve the program's problem. Memory exhaustion is the runtime's problem. The conversation is the user's asset, and it is not recoverable once summarized away. Trading irreplaceable user state for reclaimable runtime headroom is a poor exchange in general, and a surprising one when it happens without consent or notice.

It has no termination condition. The trigger is "memory is high"; the action is "compact." Nothing closes that loop by checking whether the action moved the trigger. When the memory lives anywhere other than the conversation, the condition remains true forever and the action repeats forever. That is a livelock, and it is what happened here — the guard that eventually stopped it was an unrelated 128 KB size floor, not any notion of "this isn't working."

It is most destructive exactly where continuity matters most. Long sessions accumulate the most context, are the most expensive to rebuild, and are the most likely to hit memory limits. The mechanism is therefore most aggressive precisely where the loss hurts most, and where the user is least able to notice what went missing.

Compaction is lossy, and cheaper lossless levers exist. Before discarding semantic content, a runtime under memory pressure has options that cost the user nothing: drop cached tool outputs, release parsed representations of on-disk data, stream events.jsonl rather than holding it resident, free transient buffers. Compaction destroys meaning and should be the last resort, not the first response.

It is silent and therefore indistinguishable from a bug. From the user's side, context vanished at 23% window usage with no explanation. There is no way to tell a memory-triggered compaction from a context-triggered one, so there is no way to respond correctly — and the natural conclusion is that the tool is broken.

There is no opt-out. No environment variable or setting disables memory-triggered compaction. A user who would rather the process grow, swap, or even crash — and keep a resumable transcript on disk — cannot express that preference. A crash is often less costly than a silent lossy summarization, because the on-disk history survives a crash intact.

The root cause is usually unbounded growth elsewhere. In this session the resident memory was dominated by the session's own events.jsonl (381 MB). Compacting the conversation does not shrink that file or its parsed form. Bounding that growth removes the pressure at its source; compaction only ever treated the symptom, and treated it in the wrong place.

Concretely, we would ask for: an effectiveness check and cooldown so it cannot loop; lossless reclamation attempted before lossy compaction; a clear UI signal when compaction is memory-triggered rather than context-triggered; and an opt-out for users who prefer a recoverable crash to silent context loss.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

レポートでは /.copilot/settings.json、/.copilot/session-state//events.jsonl、~/.copilot/logs が示されています。まず、logLevel を all に設定した長時間実行セッションを再現し、コンテキスト使用量、コンパクションの頻度、プロセスメモリ、ログの増加量を測定してください。完了条件には、効果のないコンパクションの繰り返しを防止すること、会話データを保持すること、過剰なログを制限またはフィルタリングすること、メモリによってトリガーされたコンパクションを明確に報告することを含める必要があります。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
linux, node.js, rust
領域
cli, devtools, observability, performance
issue の種類
バグ
難易度
5/5
見積もり時間
1週間以上
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
38/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。