CommandCodeAI / CommandCodeAI/command-code

Provider API: failed requests and no-op agent runs are billed in full ($15.56 in under 3 minutes on gpt-5.6-sol)

Aperta
#888 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Lingua principale
Nessun dato sulla lingua
Stelle
4k
Fork
350
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Summary

Provider API, model gpt-5.6-sol, plan individual-goat:
$15.5624 charged across 12 requests in 2 minutes 49 seconds (2026-09-19 05:36:24-05:39:13 local UTC+8 / 2026-09-18 21:36-21:39 UTC).

Each request carried ~305k-336k tokens of context, giving an effective rate of ~$4.02 per 1M tokens. This single 3-minute burst exceeded the 5-hour window cap ($14; window shows $17.2783 used) and consumed 44% of the $35 weekly cap.

Expected Behavior
  1. A request that returns FAILED should not be billed in full.
  2. Exceeding a 5-hour window should not be reachable by one client in under 3 minutes with no warning.
  3. A per-API-key spend cap or alert (and/or a warning threshold before the window is exhausted) would prevent silent burn.
Actual Behavior
  • The request at 05:39:13 returned FAILED and was still billed $0.2844.
  • A background agent made 10 calls on gpt-5.6-sol costing ~$7.19 with no usable output at all.
  • 4 of the 12 requests cost ~$3.00-3.31 each ($12.6744 = 81.4% of the spend); the other 8 cost $0.27-0.44. The expensive ones appear to be full cache-prefix rewrites billed at $6.25/1M.
  • Net: $15.5624 for 3,869,607 tokens (1,306,770 input + 2,562,837 cache read + 14,367 output) = $4.02 per 1M tokens.
Steps to reproduce the issue
  1. Use the Provider API (api.commandcode.ai/provider/v1) with model gpt-5.6-sol on a plan that has a 5-hour window cap.
  2. Send requests from a client that carries a large context (~300k+ tokens per call) and does not keep the prompt prefix stable - e.g. a harness that compacts context, retries, or switches models.
  3. Observe ~$1.30 average per request, then the 5-hour window being exceeded within minutes.
Command Code Version

1.56.2

Operating System

Windows

Terminal/IDE

DeepSeek Harness (dsh) via Provider API, not the CLI

Shell

PowerShell

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

Requesting a review and credit for the FAILED request ($0.2844) and the 10 no-op background-agent calls (~$7.19) - failed or empty runs should not bill at full rate.

Account context:

  • plan: individual-goat; billing period started 2026-09-18T12:02:57Z
  • official usage summary for the period: $19.2692 over 1,573 requests, 670,391,454 input tokens / 1,170,128 output tokens
  • /alpha/billing/credits at the time: fiveHour used $17.2783 / cap $14 (exceeded), weekly used $19.2547 / cap $35

Feature request that would have prevented this: expose per-request cost in the API (or a spend webhook) so clients can stop before burning a whole window. Note also that the usage API does not break out cache-write tokens, which makes client-side cost meters under-report roughly 2x for models that bill cache writes.

Happy to share the 12 request IDs from the Studio usage log.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia dal punto di ingresso di Provider API, api.commandcode.ai/provider/v1, e traccia come vengono contabilizzate le richieste fallite, le esecuzioni vuote degli agenti, i token di scrittura della cache e i limiti di spesa. Non viene indicato alcun file del repository né alcun test; il completamento dovrebbe stabilire se le chiamate fallite o no-op vengono fatturate correttamente e se viene affrontato il comportamento del limite riportato, dell'avviso o del costo per richiesta.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Ambito
ai, api, payments
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
42/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.