Long-running agent sessions exhaust event storage and appear cancelled while CLI remains active
Personne n'a encore pris cette issue.
- Langage dominant
- Shell
- Étoiles
- 11.2k
- Forks
- 1.9k
- Merge moyen
- 14 h 16 min
- PR mergées (30 j)
- 6
Description
Description
Long-running Copilot CLI project sessions that spawn many subagents can exhaust the remote session event store. Once this happens, session status and handoffs become unreliable: sessions appear inactive or cancelled even though their CLI processes remain alive and may still be executing subagents.
Environment
- Copilot CLI:
1.0.79-5 - OS: Windows
- Workload: multiple long-running project/worktree sessions using custom agents with planner, debugger, impact-analysis, coverage, and review subagents
Reproduction
-
Start a project session that performs a multi-hour workflow with frequent tool calls, planner checkpoints, and nested subagents.
-
Allow the session to accumulate a large event history.
-
Observe repeated exporter failures such as:
Failed to submit events to Mission Control session …: 409 {"code":"session_event_storage_exhausted","message":"session event storage exhausted"} remote session batch flush failed … "event_count":500 remote session exporter circuit opened after batch flush failure -
Query the session through the app/runtime status APIs.
Actual behavior
- The project session may lose its
active_session_idor appear inactive. - Background-agent status may report
Cancelledor no agents, without a cancellation reason. - Final handoffs may not arrive.
- Restarting based on this status can create a second writer for the same dirty worktree.
- The underlying CLI process can still be alive, consuming CPU, updating its log, and executing internal subagents after the UI/runtime reports no active work.
- Logs can also report cleanup of many orphaned tool calls, but do not identify a timeout, quota, or other cancellation cause.
Expected behavior
- Event storage should rotate, compact, or start a continuation stream before exhaustion.
- Exporter failure should not make an otherwise-running session unsteerable or misreport its liveness.
- Status should distinguish at least: CLI alive, agent running, agent idle, export disconnected, and task cancelled.
- Cancellation should always expose a reason such as user request, timeout, execution budget, parent teardown, or storage failure.
- The runtime should prevent or clearly warn about concurrent writers when recovery is attempted for a worktree whose original process is still active.
Impact
This makes long-running autonomous work unsafe to recover. Users cannot tell whether work has stopped, and a reasonable restart attempt can cause concurrent edits in the same worktree. Dirty changes were preserved in this case, but reliable coordination required inspecting local process IDs and logs rather than the product status APIs.
Suggested mitigation
Automatically rotate or compact session event history near the storage limit, preserve a local liveness/control channel when remote export fails, and surface a prominent event storage exhausted health state with recovery guidance.
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez par reproduire la session de projet de plusieurs heures et examinez les échecs de l’exporter, les APIs d’état du runtime ainsi que les journaux locaux des processus et des sessions décrits dans le rapport. Comparez l’état signalé avec le CLI et les sous-agents qui continuent de s’exécuter. Le travail est considéré comme terminé lorsque l’épuisement du stockage est pris en charge ou signalé clairement, que les états de liveness et les raisons d’annulation sont distinguables, et que la récupération ne peut pas créer de writers concurrents.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- shell
- Domaine
- backend-api-design, cli, observability
- Type d'issue
- Bug
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Calme
- Clarté
- Plutôt claire
- Accessibilité débutants
- 35/100