github / github/copilot-cli

Sessions become unrevivable: a stale `inuse.<pid>.lock` from a crashed host is never reclaimed on open

オープン
#4,805 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

triage
主要言語
Shell
スター
11.2k
フォーク
1.9k
平均マージ
14時間 16分
マージ済み PR(30日)
6

説明

Describe the bug

Several saved sessions became impossible to reopen ("crashed out / unrevivable"). On investigation the session data was not corrupt — the append-only events.jsonl replayed cleanly and the engine logged a successful resume. The sessions were blocked by runtime/lifecycle issues, not data loss.

The primary cause is stale lock handling. Each live session writes ~/.copilot/session-state/<id>/inuse.<pid>.lock. When the owning copilot --server process exits uncleanly (crash, force-quit, OS kill), the lock file is left behind. On the next open, the session is treated as in-use / unrevivable instead of checking whether that PID is still alive.

Scanning one local store, 14 of 184 sessions (~8%) carried a fault signature; 8 held an inuse.<pid>.lock whose PID was dead. Deleting the stale lock made each session openable again, and the app then spawned a fresh host and resumed normally — confirming the data was fine and only the lock blocked recovery.

Three related lifecycle defects compound the impact (details under Additional context):

  1. Orphaned engine host after a failed UI attach — the engine resumes but no session host registers, leaving a live server holding the lock (which then becomes defect #1 on the next attempt).
  2. MCP OAuth timeouts block session readiness — resume stalls ~5 minutes on OAuth callback timeout before the session is interactive, so users abort.
  3. Unbounded events.jsonl makes replay hang on large/old sessions (appears to overlap #4251).
Affected version

1.0.83 (copilot --server --stdio --no-auto-update)

Steps to reproduce the behavior

Primary defect (stale lock not reclaimed):

  1. Open a session, then hard-kill its copilot --server process to simulate a crash / force-quit (kill -9 <pid>).
  2. Confirm ~/.copilot/session-state/<id>/inuse.<pid>.lock remains, now pointing at a dead PID.
  3. Try to reopen the session → it does not revive.
  4. rm the stale lock file → the session reopens and resumes cleanly (a fresh host is spawned and events.jsonl replays without error).
Expected behavior

On open, if the lock's PID is dead or a zombie, the app should reclaim the session automatically rather than treating it as in-use. A liveness check combined with a PID start-time comparison avoids the classic PID-reuse race. Recoverable work should never be gated behind a leftover lock file that only manual filesystem surgery can clear.

Additional context
  • OS: macOS · shell: zsh

Defect 2 — orphaned engine host after a failed UI attach. Engine log from a single resume: the host never registers, yet the engine reports the session resumed, so the process lingers as an orphan holding the lock:

[WARNING] [rust:copilot_runtime::session::pending_request_flow] pending-request event was not delivered to the session host
          {"error":"GenericFailure, no session host is registered for session <id>"}
...
[INFO]  Resumed session: <id>        (logged repeatedly)

Suggested fix: tie the engine host's lifetime to successful UI host registration; if registration fails/times out, tear the host down and release the lock.

Defect 3 — MCP OAuth timeouts block resume (~5 min). With several MCP servers configured, resume stalls until interactive OAuth flows time out. MCP init begins at 20:27, first hard failures at 20:32:

[ERROR] [rust:mcp_engine::session_authorizer] OAuth login failed after returning its authorization URL
        {"server_name":"...","error":"OAuth callback timeout"}

During that window the session looks hung and users abort — observed as sessions whose only recent events were resume → abort (user_initiated) → shutdown, twice in a row. Suggested fix: make MCP init lazy / non-blocking; never gate session interactivity on MCP OAuth.

Defect 4 — unbounded events.jsonl. The append-only log grows without bound; worst local session reached 38.6 MB / ~13k events with 42 compaction cycles (173 re-injected system.message events ≈ 10.7 MB). Reopening replays the whole log, so large sessions hang on open even though nothing is corrupt. This looks like the same root cause as #4251; noting here for linkage. Suggested fix: snapshot-and-truncate the event log at compaction so replay cost stays bounded.

Possibly related: #4020 (session falsely "already in use by another client"), #4098 (truncated/concatenated events on resume), #4138 (resume compaction hangs), #4251 (large-session resume OOM/hang).

All four are addressable without changing the on-disk event format. Defects 1–2 turn recoverable work into apparent data loss; 3–4 make large/older sessions feel dead on open.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず、~/.copilot/session-state//inuse..lock の session-open 処理と copilot --server のライフサイクルを追跡します。強制終了のケースを再現し、その後、ロックの回収とホスト登録に関する Rust のセッションランタイムを調査します。デッドまたはゾンビ状態の PID によって再オープンがブロックされず、手動でロックを削除しなくてもセッションが再開されれば、主な不具合は解消されたことになります。OAuth と events.jsonl に関する懸念は、別のスコープとして扱う必要があります。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
rust
領域
backend, cli
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
45/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。