something seems not correct with ds v4 flash caching
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 28/100
Hướng nghiên cứu
Bắt đầu bằng việc xem xét bằng chứng từ usage dashboard được đính kèm và các request trace IDs gần đây được đề cập trong báo cáo. So sánh input bị tính phí với hành vi cache-read dự kiến và xác minh xem các request DeepSeek V4 Flash có nhận được prompt-cache hits hay không; hoàn tất có nghĩa là xác nhận nguyên nhân do routing hoặc caching, hoặc loại trừ nguyên nhân đó bằng bằng chứng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
Hi Command team,
Could you please help check whether there is an issue with DeepSeek V4 Flash caching or upstream routing on the Command servers?
My GOAT usage has increased unusually fast over the last couple of hours, and I hit the $35 weekly limit much faster than before, even though I have only been using deepseek/deepseek-v4-flash for past few hours and my workflow has not changed significantly.
One recent request from my usage dashboard shows:
- Input tokens: 357,723
- Output tokens: 120
- Charged: $0.0783
- Model:
deepseek/deepseek-v4-flash
At the current off-peak DeepSeek V4 Flash pricing, 357,723 fresh input tokens × $0.22/M is approximately $0.0787, which is almost exactly what I was charged.
This seems to suggest that nearly the entire input was billed as fresh/cache-miss input rather than cache-read input.
This is unusual for my workload. When I previously used the DeepSeek API directly for the same type of long-running agent workflow, my prompt cache hit rate was typically around 95–98%. For example, on one day I had about 758M cached input tokens versus only 13.4M uncached input tokens.
Could you please check whether there has recently been any issue with:
- DeepSeek V4 Flash prompt caching
- upstream provider/server routing or cache affinity
- cache keys being invalidated between consecutive agent requests
- any changes related to the recent DeepSeek V4 Flash / pricing update
I’ve attached screenshots of both my Command usage and my previous DeepSeek API usage for comparison.
If possible, could you also check the trace IDs from my recent requests and confirm whether those requests were actually receiving prompt-cache hits?
Thanks!
Expected Behavior
n.a.
Actual Behavior
n.a.
Steps to reproduce the issue
n.a.
Command Code Version
claude code
Operating System
macOS
Terminal/IDE
No response
Shell
No response
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
No response
- Ngôn ngữ chính
- Không có dữ liệu ngôn ngữ
- Star
- 4k
- Fork
- 350
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của CommandCodeAI/command-code
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#855 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
CommandCodeAI/command-code#841 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#655 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#608 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 70/100
CommandCodeAI/command-code#893 ·
Tất cả issue của CommandCodeAI/command-code
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
use-agent-os/agent-os#3263 ·
-
[Bug]: context-limit error parsing has no pattern for llama.cpp's "context size (N tokens)" phrasing Đang mởarea/compression area/local-models area/sessions comp/agent duplicate P2 sweeper:risk-session-state type/bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NousResearch/hermes-agent#117793 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
BasedHardware/omi#15236 · 1 bình luận ·