github / github/copilot-cli

Long-running agent sessions exhaust event storage and appear cancelled while CLI remains active

オープン
#4,467 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

area:agents area:sessions
主要言語
Shell
スター
11.2k
フォーク
1.9k
平均マージ
14時間 16分
マージ済み PR(30日)
6

説明

Description

Long-running Copilot CLI project sessions that spawn many subagents can exhaust the remote session event store. Once this happens, session status and handoffs become unreliable: sessions appear inactive or cancelled even though their CLI processes remain alive and may still be executing subagents.

Environment

  • Copilot CLI: 1.0.79-5
  • OS: Windows
  • Workload: multiple long-running project/worktree sessions using custom agents with planner, debugger, impact-analysis, coverage, and review subagents

Reproduction

  1. Start a project session that performs a multi-hour workflow with frequent tool calls, planner checkpoints, and nested subagents.

  2. Allow the session to accumulate a large event history.

  3. Observe repeated exporter failures such as:

    Failed to submit events to Mission Control session …: 409
    {"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
    remote session batch flush failed … "event_count":500
    remote session exporter circuit opened after batch flush failure
    
  4. Query the session through the app/runtime status APIs.

Actual behavior

  • The project session may lose its active_session_id or appear inactive.
  • Background-agent status may report Cancelled or no agents, without a cancellation reason.
  • Final handoffs may not arrive.
  • Restarting based on this status can create a second writer for the same dirty worktree.
  • The underlying CLI process can still be alive, consuming CPU, updating its log, and executing internal subagents after the UI/runtime reports no active work.
  • Logs can also report cleanup of many orphaned tool calls, but do not identify a timeout, quota, or other cancellation cause.

Expected behavior

  • Event storage should rotate, compact, or start a continuation stream before exhaustion.
  • Exporter failure should not make an otherwise-running session unsteerable or misreport its liveness.
  • Status should distinguish at least: CLI alive, agent running, agent idle, export disconnected, and task cancelled.
  • Cancellation should always expose a reason such as user request, timeout, execution budget, parent teardown, or storage failure.
  • The runtime should prevent or clearly warn about concurrent writers when recovery is attempted for a worktree whose original process is still active.

Impact

This makes long-running autonomous work unsafe to recover. Users cannot tell whether work has stopped, and a reasonable restart attempt can cause concurrent edits in the same worktree. Dirty changes were preserved in this case, but reliable coordination required inspecting local process IDs and logs rather than the product status APIs.

Suggested mitigation

Automatically rotate or compact session event history near the storage limit, preserve a local liveness/control channel when remote export fails, and surface a prominent event storage exhausted health state with recovery guidance.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず数時間にわたるプロジェクトセッションを再現し、レポートで説明されている exporter の失敗、runtime ステータス API、ローカルのプロセスログとセッションログを調査します。報告されたステータスを、まだ実行中の CLI およびサブエージェントと比較します。ストレージの枯渇が処理されるか明確に示され、liveness 状態とキャンセル理由を区別でき、リカバリによって同時書き込みを行う writer が作成されないことを確認できれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
shell
領域
backend-api-design, cli, observability
issue の種類
バグ
難易度
5/5
見積もり時間
1週間以上
活発さ
静か
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。