github / github/copilot-cli

Long-running agent sessions exhaust event storage and appear cancelled while CLI remains active

未关闭
#4,467 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

area:agents area:sessions
主要语言
Shell
星标
11.2k
派生
1.9k
平均合并
14 小时 16 分钟
30 天内合并 PR
6

描述

Description

Long-running Copilot CLI project sessions that spawn many subagents can exhaust the remote session event store. Once this happens, session status and handoffs become unreliable: sessions appear inactive or cancelled even though their CLI processes remain alive and may still be executing subagents.

Environment

  • Copilot CLI: 1.0.79-5
  • OS: Windows
  • Workload: multiple long-running project/worktree sessions using custom agents with planner, debugger, impact-analysis, coverage, and review subagents

Reproduction

  1. Start a project session that performs a multi-hour workflow with frequent tool calls, planner checkpoints, and nested subagents.

  2. Allow the session to accumulate a large event history.

  3. Observe repeated exporter failures such as:

    Failed to submit events to Mission Control session …: 409
    {"code":"session_event_storage_exhausted","message":"session event storage exhausted"}
    remote session batch flush failed … "event_count":500
    remote session exporter circuit opened after batch flush failure
    
  4. Query the session through the app/runtime status APIs.

Actual behavior

  • The project session may lose its active_session_id or appear inactive.
  • Background-agent status may report Cancelled or no agents, without a cancellation reason.
  • Final handoffs may not arrive.
  • Restarting based on this status can create a second writer for the same dirty worktree.
  • The underlying CLI process can still be alive, consuming CPU, updating its log, and executing internal subagents after the UI/runtime reports no active work.
  • Logs can also report cleanup of many orphaned tool calls, but do not identify a timeout, quota, or other cancellation cause.

Expected behavior

  • Event storage should rotate, compact, or start a continuation stream before exhaustion.
  • Exporter failure should not make an otherwise-running session unsteerable or misreport its liveness.
  • Status should distinguish at least: CLI alive, agent running, agent idle, export disconnected, and task cancelled.
  • Cancellation should always expose a reason such as user request, timeout, execution budget, parent teardown, or storage failure.
  • The runtime should prevent or clearly warn about concurrent writers when recovery is attempted for a worktree whose original process is still active.

Impact

This makes long-running autonomous work unsafe to recover. Users cannot tell whether work has stopped, and a reasonable restart attempt can cause concurrent edits in the same worktree. Dirty changes were preserved in this case, but reliable coordination required inspecting local process IDs and logs rather than the product status APIs.

Suggested mitigation

Automatically rotate or compact session event history near the storage limit, preserve a local liveness/control channel when remote export fails, and surface a prominent event storage exhausted health state with recovery guidance.

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

首先重现持续数小时的项目会话,并检查报告中描述的 exporter 失败、runtime 状态 API,以及本地进程和会话日志。将报告的状态与仍在运行的 CLI 和子代理进行比较。当存储耗尽得到处理或被清晰地报告、liveness 状态和取消原因可以区分,并且恢复过程无法创建并发写入者时,即视为完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
shell
领域
backend-api-design, cli, observability
Issue 类型
缺陷
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。