CommandCodeAI / CommandCodeAI/command-code

something seems not correct with ds v4 flash caching

Aberta
#702 1 comentário 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

Linguagem predominante
Sem dados de linguagem
Estrelas
4k
Forks
350
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

Summary

Hi Command team,

Could you please help check whether there is an issue with DeepSeek V4 Flash caching or upstream routing on the Command servers?

My GOAT usage has increased unusually fast over the last couple of hours, and I hit the $35 weekly limit much faster than before, even though I have only been using deepseek/deepseek-v4-flash for past few hours and my workflow has not changed significantly.

One recent request from my usage dashboard shows:

  • Input tokens: 357,723
  • Output tokens: 120
  • Charged: $0.0783
  • Model: deepseek/deepseek-v4-flash

At the current off-peak DeepSeek V4 Flash pricing, 357,723 fresh input tokens × $0.22/M is approximately $0.0787, which is almost exactly what I was charged.

This seems to suggest that nearly the entire input was billed as fresh/cache-miss input rather than cache-read input.

This is unusual for my workload. When I previously used the DeepSeek API directly for the same type of long-running agent workflow, my prompt cache hit rate was typically around 95–98%. For example, on one day I had about 758M cached input tokens versus only 13.4M uncached input tokens.

Could you please check whether there has recently been any issue with:

  • DeepSeek V4 Flash prompt caching
  • upstream provider/server routing or cache affinity
  • cache keys being invalidated between consecutive agent requests
  • any changes related to the recent DeepSeek V4 Flash / pricing update

I’ve attached screenshots of both my Command usage and my previous DeepSeek API usage for comparison.

If possible, could you also check the trace IDs from my recent requests and confirm whether those requests were actually receiving prompt-cache hits?

Thanks!

Image
Expected Behavior

n.a.

Actual Behavior

n.a.

Steps to reproduce the issue

n.a.

Command Code Version

claude code

Operating System

macOS

Terminal/IDE

No response

Shell

No response

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

No response

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece revisando as evidências anexadas do painel de uso e os request trace IDs recentes mencionados no relatório. Compare a entrada cobrada com o comportamento esperado de cache-read e verifique se as solicitações do DeepSeek V4 Flash estão recebendo prompt-cache hits; o trabalho estará concluído quando a causa de routing ou caching for confirmada, ou descartada com evidências.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Domínio
ai, backend
Tipo de issue
Bug
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Ativa
Clareza
Precisa de esclarecimento
Facilidade para iniciantes
28/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.