Sessions become unrevivable: a stale `inuse.<pid>.lock` from a crashed host is never reclaimed on open
Nobody has claimed this yet.
- Dominant language
- Shell
- Stars
- 11.2k
- Forks
- 1.9k
- Avg merge
- 14h 16m
- Merged PRs (30d)
- 6
Description
Describe the bug
Several saved sessions became impossible to reopen ("crashed out / unrevivable"). On investigation the session data was not corrupt — the append-only events.jsonl replayed cleanly and the engine logged a successful resume. The sessions were blocked by runtime/lifecycle issues, not data loss.
The primary cause is stale lock handling. Each live session writes ~/.copilot/session-state/<id>/inuse.<pid>.lock. When the owning copilot --server process exits uncleanly (crash, force-quit, OS kill), the lock file is left behind. On the next open, the session is treated as in-use / unrevivable instead of checking whether that PID is still alive.
Scanning one local store, 14 of 184 sessions (~8%) carried a fault signature; 8 held an inuse.<pid>.lock whose PID was dead. Deleting the stale lock made each session openable again, and the app then spawned a fresh host and resumed normally — confirming the data was fine and only the lock blocked recovery.
Three related lifecycle defects compound the impact (details under Additional context):
- Orphaned engine host after a failed UI attach — the engine resumes but no session host registers, leaving a live server holding the lock (which then becomes defect #1 on the next attempt).
- MCP OAuth timeouts block session readiness — resume stalls ~5 minutes on
OAuth callback timeoutbefore the session is interactive, so users abort. - Unbounded
events.jsonlmakes replay hang on large/old sessions (appears to overlap #4251).
Affected version
1.0.83 (copilot --server --stdio --no-auto-update)
Steps to reproduce the behavior
Primary defect (stale lock not reclaimed):
- Open a session, then hard-kill its
copilot --serverprocess to simulate a crash / force-quit (kill -9 <pid>). - Confirm
~/.copilot/session-state/<id>/inuse.<pid>.lockremains, now pointing at a dead PID. - Try to reopen the session → it does not revive.
rmthe stale lock file → the session reopens and resumes cleanly (a fresh host is spawned andevents.jsonlreplays without error).
Expected behavior
On open, if the lock's PID is dead or a zombie, the app should reclaim the session automatically rather than treating it as in-use. A liveness check combined with a PID start-time comparison avoids the classic PID-reuse race. Recoverable work should never be gated behind a leftover lock file that only manual filesystem surgery can clear.
Additional context
- OS: macOS · shell: zsh
Defect 2 — orphaned engine host after a failed UI attach. Engine log from a single resume: the host never registers, yet the engine reports the session resumed, so the process lingers as an orphan holding the lock:
[WARNING] [rust:copilot_runtime::session::pending_request_flow] pending-request event was not delivered to the session host
{"error":"GenericFailure, no session host is registered for session <id>"}
...
[INFO] Resumed session: <id> (logged repeatedly)
Suggested fix: tie the engine host's lifetime to successful UI host registration; if registration fails/times out, tear the host down and release the lock.
Defect 3 — MCP OAuth timeouts block resume (~5 min). With several MCP servers configured, resume stalls until interactive OAuth flows time out. MCP init begins at 20:27, first hard failures at 20:32:
[ERROR] [rust:mcp_engine::session_authorizer] OAuth login failed after returning its authorization URL
{"server_name":"...","error":"OAuth callback timeout"}
During that window the session looks hung and users abort — observed as sessions whose only recent events were resume → abort (user_initiated) → shutdown, twice in a row. Suggested fix: make MCP init lazy / non-blocking; never gate session interactivity on MCP OAuth.
Defect 4 — unbounded events.jsonl. The append-only log grows without bound; worst local session reached 38.6 MB / ~13k events with 42 compaction cycles (173 re-injected system.message events ≈ 10.7 MB). Reopening replays the whole log, so large sessions hang on open even though nothing is corrupt. This looks like the same root cause as #4251; noting here for linkage. Suggested fix: snapshot-and-truncate the event log at compaction so replay cost stays bounded.
Possibly related: #4020 (session falsely "already in use by another client"), #4098 (truncated/concatenated events on resume), #4138 (resume compaction hangs), #4251 (large-session resume OOM/hang).
All four are addressable without changing the on-disk event format. Defects 1–2 turn recoverable work into apparent data loss; 3–4 make large/older sessions feel dead on open.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing session-open handling for ~/.copilot/session-state//inuse..lock and the copilot --server lifecycle. Reproduce the hard-kill case, then inspect the Rust session runtime around lock reclamation and host registration; the primary defect is done when a dead or zombie PID no longer blocks reopening and the session resumes without manual lock removal. The OAuth and events.jsonl concerns should be treated as separate scope.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- backend, cli
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100