something seems not correct with ds v4 flash caching
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 28/100
Rechercherichtung
Beginnen Sie mit der Überprüfung der beigefügten Belege aus dem Nutzungsdashboard und der im Bericht genannten aktuellen request trace IDs. Vergleichen Sie die berechnete Eingabe mit dem erwarteten cache-read-Verhalten und verifizieren Sie, ob DeepSeek V4 Flash requests prompt-cache hits erhalten; abgeschlossen bedeutet, die Routing- oder Caching-Ursache zu bestätigen oder sie anhand von Belegen auszuschließen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Summary
Hi Command team,
Could you please help check whether there is an issue with DeepSeek V4 Flash caching or upstream routing on the Command servers?
My GOAT usage has increased unusually fast over the last couple of hours, and I hit the $35 weekly limit much faster than before, even though I have only been using deepseek/deepseek-v4-flash for past few hours and my workflow has not changed significantly.
One recent request from my usage dashboard shows:
- Input tokens: 357,723
- Output tokens: 120
- Charged: $0.0783
- Model:
deepseek/deepseek-v4-flash
At the current off-peak DeepSeek V4 Flash pricing, 357,723 fresh input tokens × $0.22/M is approximately $0.0787, which is almost exactly what I was charged.
This seems to suggest that nearly the entire input was billed as fresh/cache-miss input rather than cache-read input.
This is unusual for my workload. When I previously used the DeepSeek API directly for the same type of long-running agent workflow, my prompt cache hit rate was typically around 95–98%. For example, on one day I had about 758M cached input tokens versus only 13.4M uncached input tokens.
Could you please check whether there has recently been any issue with:
- DeepSeek V4 Flash prompt caching
- upstream provider/server routing or cache affinity
- cache keys being invalidated between consecutive agent requests
- any changes related to the recent DeepSeek V4 Flash / pricing update
I’ve attached screenshots of both my Command usage and my previous DeepSeek API usage for comparison.
If possible, could you also check the trace IDs from my recent requests and confirm whether those requests were actually receiving prompt-cache hits?
Thanks!
Expected Behavior
n.a.
Actual Behavior
n.a.
Steps to reproduce the issue
n.a.
Command Code Version
claude code
Operating System
macOS
Terminal/IDE
No response
Shell
No response
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
No response
- Vorherrschende Sprache
- Keine Sprachdaten
- Sterne
- 4k
- Forks
- 350
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus CommandCodeAI/command-code
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#855 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
CommandCodeAI/command-code#841 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#655 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#608 ·
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 70/100
CommandCodeAI/command-code#893 ·
Alle Issues in CommandCodeAI/command-code
Ähnliche Issues
-
enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
JuliusBrussee/caveman#1102 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
use-agent-os/agent-os#3263 ·
-
[Bug]: context-limit error parsing has no pattern for llama.cpp's "context size (N tokens)" phrasing Offenarea/compression area/local-models area/sessions comp/agent duplicate P2 sweeper:risk-session-state type/bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
NousResearch/hermes-agent#117793 · 1 Kommentar ·
-
possible bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Mintplex-Labs/anything-llm#6415 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100