CommandCodeAI / CommandCodeAI/command-code
something seems not correct with ds v4 flash caching
還沒有人認領這個 Issue。
- 主要語言
- 沒有語言資料
- 星號
- 4k
- 分支
- 350
- PR 合併指標
- 30 天內沒有已合併 PR
描述
Summary
Hi Command team,
Could you please help check whether there is an issue with DeepSeek V4 Flash caching or upstream routing on the Command servers?
My GOAT usage has increased unusually fast over the last couple of hours, and I hit the $35 weekly limit much faster than before, even though I have only been using deepseek/deepseek-v4-flash for past few hours and my workflow has not changed significantly.
One recent request from my usage dashboard shows:
- Input tokens: 357,723
- Output tokens: 120
- Charged: $0.0783
- Model:
deepseek/deepseek-v4-flash
At the current off-peak DeepSeek V4 Flash pricing, 357,723 fresh input tokens × $0.22/M is approximately $0.0787, which is almost exactly what I was charged.
This seems to suggest that nearly the entire input was billed as fresh/cache-miss input rather than cache-read input.
This is unusual for my workload. When I previously used the DeepSeek API directly for the same type of long-running agent workflow, my prompt cache hit rate was typically around 95–98%. For example, on one day I had about 758M cached input tokens versus only 13.4M uncached input tokens.
Could you please check whether there has recently been any issue with:
- DeepSeek V4 Flash prompt caching
- upstream provider/server routing or cache affinity
- cache keys being invalidated between consecutive agent requests
- any changes related to the recent DeepSeek V4 Flash / pricing update
I’ve attached screenshots of both my Command usage and my previous DeepSeek API usage for comparison.
If possible, could you also check the trace IDs from my recent requests and confirm whether those requests were actually receiving prompt-cache hits?
Thanks!
Expected Behavior
n.a.
Actual Behavior
n.a.
Steps to reproduce the issue
n.a.
Command Code Version
claude code
Operating System
macOS
Terminal/IDE
No response
Shell
No response
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
No response
貢獻指南
這個儲存庫沒有索引到貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
首先檢查附加的使用情況儀表板證據,以及報告中提到的近期 request trace IDs。將計費的輸入與預期的 cache-read 行為進行比較,並驗證 DeepSeek V4 Flash requests 是否獲得了 prompt-cache hits;完成的標準是確認 routing 或 caching 原因,或用證據排除該原因。
由索引模型根據 Issue 內容生成。
評估
- 領域
- ai, backend
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 活躍
- 描述清晰度
- 需要釐清
- 新手友好度
- 28/100