github / github/copilot-cli

Bug: Copilot CLI 1.0.82 BYOK silently disables prompt caching (~5x cost)

オープン
#4,720 コメント 0 件 リアクション 1 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

area:models area:networking
主要言語
Shell
スター
11.2k
フォーク
1.9k
平均マージ
14時間 16分
マージ済み PR(30日)
6

説明

Summary

Copilot CLI 1.0.82 in BYOK mode sends chat requests that carry no prompt-cache declaration, so prompt caching is effectively disabled for the whole session. Every turn re-sends the entire growing context at full price — provider usage shows cached_tokens=0 and cache_write/cache_creation_tokens=0 on every request across the session, even though the prefix is stable.

Net effect: ~5x cost for normal multi-turn BYOK sessions.

Setup

  • Copilot CLI 1.0.82, BYOK mode (custom OpenAI-compatible /chat/completions endpoint serving an Anthropic model, claude-sonnet-5)
  • Compare baseline: same machine, same endpoint/model, 1.0.80 BYOK → prompt caching works (~97% hit)
  • Also tested 1.0.82 in GitHub-subscription mode on the same machine → caching works, so the build is cache-capable

Expected

Once a conversation prefix stabilizes, the provider should report cached_tokens > 0 and small cache_creation_tokens for the new tail (that's what 1.0.80 BYOK and subscription mode report).

Actual

Every request reports zero cache metrics:

"prompt_tokens_details": {
  "text_tokens": 146022,
  "cached_tokens": 0,
  "cache_write_tokens": 0,
  "cache_creation_tokens": 0,
  "cache_creation_token_details": { "ephemeral_5m_input_tokens": 0 }
},
"cache_read_input_tokens": 0

Over a 150-request session, text_tokens climbed monotonically 118,966 → 285,085 while cached_tokens and cache_write/creation_tokens stayed 0 every time. The zero cache-write/creation is the strong signal: even a cold miss should write the prefix so the next turn can hit; a persistent 0 write means the request never asks for a cache at all.

Evidence

Metric 1.0.82 BYOK (broken) 1.0.80 BYOK (baseline)
Requests 150 320
Input tokens 30,075,042 100,762,401
Cached tokens 0 98,424,636
Cache hit rate 0% 97.7%
Cost $61.21 $28.92

With normal caching, 30.1M input ≈ $11.4; at 0% it costs ≈ $60.2 → ~$49 (~5x) wasted in a single session.

Root-cause hypothesis

The 1.0.82 BYOK request builder omits the prompt-cache declaration (cache_control / ephemeral-cache policy) that subscription mode injects. Because BYOK calls the user's endpoint directly, the missing declaration is only visible to BYOK users.

Severity

High — silent runaway cost on BYOK setups; no CLI error, only a much larger bill. Requires monitoring provider cache hit-rate to detect.

Suggested fix

Make the BYOK path send the same prompt-cache declaration used by subscription mode.

Repro note

I cannot reproduce right now (downgraded to 1.0.80 where caching works). Reproducing requires running 1.0.82 in BYOK against an Anthropic-model endpoint and reading provider cache accounting. Happy to help test a newer build backported if there's a fix candidate.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

BYOK request builder を特定し、その request の構築を subscription path および issue に記載されている 1.0.80 の動作と比較します。prompt-cache declaration が存在するかを確認し、その後、Anthropic-model endpoint に対して multi-turn BYOK セッションを実行して、cache-write と cache-read のメトリクスがゼロではなくなっていることを検証します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
shell
領域
api, cli
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
55/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。