Bad default: engine falls back to 128K token budget for model
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Mô tả
Summary
When a routed model has no capability limits installed (or reports a zero context window), the agent engine silently falls back to a hardcoded 128,000-token prompt budget and drives context compaction against it. For large-context models (e.g. 1M-token Anthropic models) whose ids don't match a host-provided capabilities entry, this makes compaction fire ~8x too early (at ~102K tokens instead of ~800K+), causing avoidable summarization churn, latency, token cost, and quality loss.
We hit this integrating the SDK/engine into a host app. We can work around it host-side by always installing capabilities, but the engine's default is not sane and the fallback operator makes it worse.
Where it lives
In the bundled engine (@github/copilot, observed in 1.0.63 and 1.0.73; app.js, minified). The relevant compiled forms:
// module init
sen = .8, aen = .95, Nk = 128e3; // Nk == 128000
class LM { static DEFAULT_TOKEN_LIMIT = Nk; /* ... */ }
// CompactionProcessor.preRequest
s = r.capabilities?.limits?.max_prompt_tokens
|| r.capabilities?.limits?.max_context_window_tokens
|| Nk; // <-- falsy fallback to 128000
u = promptTokens + toolTokens;
d = u / s; // utilization
// compaction triggers when d >= 0.8 (background) / 0.95 (buffer exhaustion)
The same ... || Nk fallback is repeated in the session.usage_info emit, contextInfo, and getTokenLimits paths.
Two distinct problems
-
Falsy fallback (
||) instead of nullish (??).
A model that reportsmax_context_window_tokens: 0(a legitimately "unknown" signal) collapses to128000rather than being treated as unknown. This is compounded upstream: the@github/copilot-sdkclient'smodels.listhandler backfills missing limits with{ max_context_window_tokens: 0 }, so an un-capped model arrives at the engine with0and|| Nkturns it into 128K.0should not be coerced to the default via a truthiness check. -
No model-aware default.
Every un-registered model — including known large-window models — gets the same 128K budget. There is no per-family/default table and no way to distinguish "small model, 128K is right" from "1M-window model, 128K is catastrophically low."
Impact
- Large-context models compact at ~0.8 * 128000 ≈ 102K tokens regardless of their true window.
- Symptoms: premature/repeated context compaction, extra summarization round-trips, higher token spend and latency, degraded answer quality on long tasks.
- Silent: nothing in the default event payload surfaces the effective window that was used, so the "capped at 128K" cause has to be inferred from the model id. (We had to add host-side telemetry — reconstructing
max_prompt_tokens ?? max_context_window_tokens ?? 128000— to see it.)
Steps to reproduce
- Create a session with a large-context model whose id is not matched by any host-installed capabilities entry (so no
limitsreach the engine, or they arrive asmax_context_window_tokens: 0). - Send a turn whose prompt+tool tokens exceed ~102K but are well under the model's real window (e.g. 300K on a 1M-window model).
- Observe
session.compaction_startfiring even though the real window is nowhere near exhausted.
Expected
0/ missing limits should be treated as "unknown," not coerced to 128K (use??, or validate> 0).- Provide a sane, model-aware default (or at minimum a host-configurable default budget) so large-window models are not capped at 128K.
- Surface the effective token limit the engine used in the compaction /
usage_infotelemetry so the applied budget is observable without host-side reconstruction.
Suggested fixes
- Change the fallback chain from
a || b || Nkto nullish/> 0validation so a real0isn't silently replaced. - Add a model-aware default table, or accept an explicit host-supplied default token budget on session config.
- Emit the resolved effective prompt-token limit alongside compaction and
session.usage_infoevents.
Environment
- Engine:
@github/copilot1.0.63(also reproduced against1.0.73via@github/copilot-next). - SDK:
@github/copilot-sdk1.0.0(zero-window backfill also present in1.0.7).
Related
The SDK-side zero-window backfill (@github/copilot-sdk models.list handler) contributes to problem (1). Happy to cross-file there if the SDK is the preferred owner for that half.
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu với app.js đi kèm và lần theo CompactionProcessor.preRequest, fallback lặp lại trong session.usage_info, contextInfo và getTokenLimits, cùng với handler models.list của SDK. Xác minh hành vi khi giới hạn bằng không hoặc bị thiếu bằng các bước tái hiện, sau đó xác định cách giới hạn token prompt hiệu dụng cần được phân giải và cung cấp trong compaction cũng như telemetry sử dụng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- javascript
- Lĩnh vực
- cli, observability
- Loại issue
- Lỗi
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức độ hoạt động
- Ít trao đổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 35/100