github / github/copilot-cli

Parallel explore subagent fan-out dies to per-model 429s: explore's default model is the only rate-limited one, no backoff, no auto model switch despite eligibleForAutoSwitch

未关闭
#4,416 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

area:agents area:models
主要语言
Shell
星标
11.2k
派生
1.9k
平均合并
14 小时 16 分钟
30 天内合并 PR
6

描述

Describe the bug

Launching many subagents in parallel via the task tool concentrates all their model calls on one model bucket — explore agents all default to the same lightweight model (currently claude-haiku-4.5). That model appears to have a much tighter per-model burst limit than any other model, so a 16-agent explore fan-out hits HTTP 429 within ~20 seconds. Every subagent then fails with repeated integration_rate_limited errors and completes with empty output, while the parent session (on a different model) continues unaffected.

Three compounding problems:

  1. The built-in explore agent defaults to the only model that rate-limits under fan-out. A survey of 1,089 local CLI sessions (~2.7 GB of events.jsonl) found exactly 5 sessions that ever logged errorType: rate_limit — every significant incident was a 16x claude-haiku-4.5 explore fan-out. Other models never triggered it despite far heavier use: gpt-5.6-sol (19k assistant messages, 201 sessions), claude-fable-5 (14k messages, including a clean 16-way fan-out re-run of the exact workload that failed on haiku), claude-sonnet-5 (9k), claude-opus-5 (7k), gpt-5.6-terra (5k) — zero incidents.
  2. No backoff. The agentic loop retried into the same window at ~2 failing requests/second: 182 integration_rate_limited errors logged in 87 seconds in one incident.
  3. No model fallback. The 429 response carries "eligibleForAutoSwitch": true, but subagents never switch — they just die and return empty results, silently wasting the whole fan-out.

Sample error from events.jsonl:

{"errorType":"rate_limit",
 "message":"You've hit the rate limit for this model. Please switch models or wait for your limit to reset in under a minute. Learn More (https://docs.github.com/copilot/concepts/rate-limits). (Request ID: C33A:2A98BA:2871A5:13FF8A5:6A7702DD)",
 "statusCode":429,
 "errorCode":"integration_rate_limited",
 "eligibleForAutoSwitch":true}

Observed concurrency data for the explore default model (same account, 4 incidents across 3 days):

  • 16 parallel explore agents -> first 429 in ~20 s, all 16 stall and return empty
  • 8 parallel -> clean
  • ~15 launched sequentially over 10 minutes with overlap -> 2 transient errors, work completed
  • 16 parallel on claude-fable-5 (model override) -> zero errors
Affected version

1.0.79-5 (Windows x64); incidents also observed on earlier 1.0.7x builds.

Steps to reproduce the behavior
  1. In a session, use the task tool to launch ~16 explore subagents simultaneously (all default to the lightweight model).
  2. Within ~20 seconds, subagent transcripts fill with "Limit reached — Resets in under a minute"; agents go idle after emitting only setup text and return empty results.
  3. Re-run the identical fan-out with a model override to a larger model — it completes cleanly.
Expected behavior

Any (ideally all) of:

  • Subagents honor eligibleForAutoSwitch and fall back to another model instead of dying (related: #2840).
  • The agentic loop backs off per the reset window instead of retrying ~2x/s into the same limited minute (related: #2760).
  • The CLI throttles subagent fan-out concurrency per model client-side, since it knows how many concurrent loops it is aiming at one bucket (related: #2545).
  • The built-in explore agent''s default model either gets burst headroom matching the "fast, lightweight, fan out in parallel" positioning, or the CLI spreads large fan-outs across multiple eligible lightweight models.
Additional context
  • The failure is invisible from the parent''s perspective until results come back empty: subagent.completed events fire normally with ~200-byte payloads.
  • Nudging stalled subagents (write_agent) during the limited minute makes it worse — each nudge adds more 429s to the same window.
  • Request IDs from multiple incidents available on request.

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

首先使用 task tool 通过 16-agent explore fan-out 复现故障,然后比较 default model 和 model override 的 events.jsonl。Done 应防止重复的 429 retries,并确保 subagents 通过 backoff、throttling 或 model switching 返回有用的结果;payload 未指定任何实现文件或测试。

由索引模型根据 Issue 内容生成。

评估

领域
ai-infra-agents, cli, performance
Issue 类型
缺陷
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。