anthropics / anthropics/claude-code

[FEATURE] Per-model prompt-cache TTL — extend promptCacheTtl to a per-model map

Aperta
#89,488 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
api:bedrock area:cost enhancement
Lingua principale
Python
Stelle
145k
Fork
23.1k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

### Preflight Checklist

- [x] I have searched [existing requests](https://github.com/anthropics/claude-code/issues?q=is%3Aissue%20label%3Aenhancement) and this feature hasn't been requested yet
- [x] This is a single feature request (not multiple features)

### Problem Statement

We run Claude Code for an enterprise team on Amazon Bedrock (token billing, self-hosted gateway).
Several developers exhausted their monthly budget in ~10 days. A representative session's cost split:

| Component | Tokens | Cost | Share |
|-------------|--------|-------|-------|
| Input | 59.6k | $0.30 | 5% |
| Output | 65.7k | $1.64 | 27% |
| Cache read | 3.1M | $1.55 | 26% |
| Cache write | 408.4k | $2.55 | 42% |
| **Total** | | $6.04 | |

Caching itself works well (each write is read back ~7.6x on average). The problem is TTL fit:
interactive sessions routinely have >5-minute think-time gaps, so the 5m cache (the default on
Bedrock) expires and unchanged context is rewritten at 1.25x the input rate.

`promptCacheTtl` / `subagentPromptCacheTtl` (v2.1.242+) split TTL by request kind, which helps.
But the right TTL also depends on the *model*, because the write premium is priced per model:
on Opus 4.8, a 5m write is $6.25/MTok and a 1h write is $10/MTok against a $0.50/MTok read.
The break-even between "pay 2x once" and "pay 1.25x on every expiry" lands differently for an
expensive model carrying long interactive sessions than for a cheap model doing short-lived
background work — and the main/subagent split can't express that when the same bucket runs
different models (mid-session /model switches, background tasks on cheaper models, or a gateway
routing different user groups to different models).

A single global value forces a compromise: 1h everywhere overpays on short-lived work; 5m
everywhere burns full-context rewrites on interactive sessions.

### Proposed Solution

Let the existing settings accept a per-model map, keeping the current string form as shorthand:

{
"promptCacheTtl": {
"default": "1h",
"claude-haiku-4-5": "5m"
},
"subagentPromptCacheTtl": "5m"
}

Longer term we'd prefer the platform to infer TTL from workload shape (short-lived subagent
requests at 5m, interactive sessions at 1h) — #56307 asked for that and was closed not-planned,
so per-model configuration is the practical next step operators can act on.

Related: #74075 (shipped the configurable TTL), #47425 (intermediate TTL tiers — complementary),
#62217 / #60316 (superseded by the shipped settings).

### Alternative Solutions

_No response_

### Priority

High - Significant impact on productivity

### Feature Category

Performance and speed

### Use Case Example

_No response_

### Additional Context

_No response_

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

The issue names the existing promptCacheTtl and subagentPromptCacheTtl settings but no files, tests, or entry points. Start by locating their configuration parsing and model-selection paths. Done means the settings accept a per-model map with a default, preserve the current string form, and apply the selected TTL correctly across main and subagent requests.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
aws
Ambito
ai, cli, cloud, performance
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
45/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.