github / github/copilot-cli

Bad default: engine falls back to 128K token budget for model

未关闭
#4,310 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area:context-memory area:models
主要语言
Shell
星标
11.2k
派生
1.9k
平均合并
14 小时 16 分钟
30 天内合并 PR
6

描述

## Summary

When a routed model has no capability limits installed (or reports a zero context window), the agent engine silently falls back to a hardcoded **128,000**-token prompt budget and drives context compaction against it. For large-context models (e.g. 1M-token Anthropic models) whose ids don't match a host-provided capabilities entry, this makes compaction fire ~8x too early (at ~102K tokens instead of ~800K+), causing avoidable summarization churn, latency, token cost, and quality loss.

We hit this integrating the SDK/engine into a host app. We can work around it host-side by always installing capabilities, but the engine's default is not sane and the fallback operator makes it worse.

## Where it lives

In the bundled engine (`@github/copilot`, observed in `1.0.63` and `1.0.73`; `app.js`, minified). The relevant compiled forms:

```js
// module init
sen = .8, aen = .95, Nk = 128e3; // Nk == 128000
class LM { static DEFAULT_TOKEN_LIMIT = Nk; /* ... */ }

// CompactionProcessor.preRequest
s = r.capabilities?.limits?.max_prompt_tokens
|| r.capabilities?.limits?.max_context_window_tokens
|| Nk; // <-- falsy fallback to 128000
u = promptTokens + toolTokens;
d = u / s; // utilization
// compaction triggers when d >= 0.8 (background) / 0.95 (buffer exhaustion)
```

The same `... || Nk` fallback is repeated in the `session.usage_info` emit, `contextInfo`, and `getTokenLimits` paths.

## Two distinct problems

1. **Falsy fallback (`||`) instead of nullish (`??`).**
A model that reports `max_context_window_tokens: 0` (a legitimately "unknown" signal) collapses to `128000` rather than being treated as unknown. This is compounded upstream: the `@github/copilot-sdk` client's `models.list` handler backfills missing limits with `{ max_context_window_tokens: 0 }`, so an un-capped model arrives at the engine with `0` and `|| Nk` turns it into 128K. `0` should not be coerced to the default via a truthiness check.

2. **No model-aware default.**
Every un-registered model — including known large-window models — gets the same 128K budget. There is no per-family/default table and no way to distinguish "small model, 128K is right" from "1M-window model, 128K is catastrophically low."

## Impact

- Large-context models compact at ~0.8 * 128000 ≈ 102K tokens regardless of their true window.
- Symptoms: premature/repeated context compaction, extra summarization round-trips, higher token spend and latency, degraded answer quality on long tasks.
- Silent: nothing in the default event payload surfaces the *effective* window that was used, so the "capped at 128K" cause has to be inferred from the model id. (We had to add host-side telemetry — reconstructing `max_prompt_tokens ?? max_context_window_tokens ?? 128000` — to see it.)

## Steps to reproduce

1. Create a session with a large-context model whose id is **not** matched by any host-installed capabilities entry (so no `limits` reach the engine, or they arrive as `max_context_window_tokens: 0`).
2. Send a turn whose prompt+tool tokens exceed ~102K but are well under the model's real window (e.g. 300K on a 1M-window model).
3. Observe `session.compaction_start` firing even though the real window is nowhere near exhausted.

## Expected

- `0` / missing limits should be treated as "unknown," not coerced to 128K (use `??`, or validate `> 0`).
- Provide a sane, model-aware default (or at minimum a host-configurable default budget) so large-window models are not capped at 128K.
- Surface the effective token limit the engine used in the compaction / `usage_info` telemetry so the applied budget is observable without host-side reconstruction.

## Suggested fixes

- Change the fallback chain from `a || b || Nk` to nullish/`> 0` validation so a real `0` isn't silently replaced.
- Add a model-aware default table, or accept an explicit host-supplied default token budget on session config.
- Emit the resolved effective prompt-token limit alongside compaction and `session.usage_info` events.

## Environment

- Engine: `@github/copilot` `1.0.63` (also reproduced against `1.0.73` via `@github/copilot-next`).
- SDK: `@github/copilot-sdk` `1.0.0` (zero-window backfill also present in `1.0.7`).

## Related

The SDK-side zero-window backfill (`@github/copilot-sdk` `models.list` handler) contributes to problem (1). Happy to cross-file there if the SDK is the preferred owner for that half.

贡献指南

打开贡献指南

调研方向

从随附的 app.js 开始,跟踪 CompactionProcessor.preRequest、session.usage_info、contextInfo 和 getTokenLimits 中重复出现的 fallback,以及 SDK 的 models.list handler。根据复现步骤验证限制为零或缺失时的行为,然后定义应如何解析有效的 prompt token 限制,并在 compaction 和使用遥测中公开该限制。

由索引模型根据 Issue 内容生成。

评估

技术栈
javascript
领域
cli, observability
Issue 类型
缺陷
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。