Requesting this stays open, with the emphasis moved.
- 主要語言
- Shell
- 星號
- 11.2k
- 分支
- 1.9k
- 平均合併
- 14 小時 16 分鐘
- 30 天內合併 PR
- 6
描述
Requesting this stays open, with the emphasis moved.
Our own situation is resolved: reducing the number of configured MCP servers restored sub-agents, and we have a workaround we can live with. That was a configuration problem on our side, and if the report were only about our environment it could be closed.
**The configuration was never the bug. The bug is that hitting the limit is undetectable.**
When a sub-agent fails this way, all of the following are true at once:
- The `task` tool returns success with an empty result. It does not error.
- The parent agent cannot distinguish "the sub-agent had nothing to say" from "the sub-agent never ran," so it may reasonably continue and act on nothing.
- Nothing is written to the session log. There are zero `[ERROR]` or `[WARN]` entries in the window where this happens; the only errors present were unrelated MCP startup failures from about a minute earlier.
- Nothing renders in the UI. There is no banner, badge, or status to indicate the agent was cut off rather than quiet.
- No warning is given when crossing the threshold. Adding one more MCP server silently disables every full-tool sub-agent, with no indication that it happened or which change caused it.
The practical cost of that combination: an entire code review returned empty four times, across three different models and both sync and background modes, before the variable was isolated. Each attempt looked like a model or agent-type problem, because those were the only things visibly changing. Considerable time went into an incorrect root cause before a controlled test pointed at tool-schema volume.
This is also likely invisible in your own telemetry. If nothing is logged or errored, these failures presumably register as successful task invocations that happened to produce no output, so the real frequency across users could be much higher than the issue count suggests.
Concrete asks, roughly in value order:
1. **Fail loudly.** If the sub-agent's context cannot be assembled within budget, return an error rather than an empty success. Even a generic message beats silence.
2. **Log it.** A single line naming the cause and the budget would have reduced this from hours to minutes.
3. **Surface it in the UI.** Distinguish "agent returned nothing" from "agent could not run."
4. **Warn at configuration time.** When adding an MCP server pushes the projected tool-schema size past the limit, say so at startup, while the user still has the context to connect cause and effect.
Items 1-3 are useful regardless of how the underlying budget problem in #3542 is resolved, because the silent-failure mode will otherwise reappear behind any future limit.
_Originally posted by @ChrisMcKee1 in https://github.com/github/copilot-cli/issues/4293#issuecomment-5122291931_
貢獻指南
研究方向
首先追蹤任務工具的子代理執行路徑,以及 MCP 伺服器的設定或啟動路徑。使用 issue #3542 作為底層預算脈絡,檢查空結果、工作階段記錄、UI 狀態和閾值警告是如何處理的。完成標準是:限制導致的失敗能夠與空成功區分開來,能夠被記錄、呈現在 UI 中,並在設定期間發出警告。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- shell
- 領域
- cli, observability
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 48/100