Provider API: failed requests and no-op agent runs are billed in full ($15.56 in under 3 minutes on gpt-5.6-sol)
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 42/100
Rechercherichtung
Beginne am Provider API-Einstiegspunkt api.commandcode.ai/provider/v1 und verfolge, wie fehlgeschlagene Anfragen, leere Agent-Ausführungen, Cache-Schreib-Tokens und Ausgabenlimits erfasst werden. Es wird keine Repository-Datei und kein Test genannt; die Erledigung sollte klären, ob fehlgeschlagene oder No-op-Aufrufe korrekt abgerechnet werden und ob das Verhalten des gemeldeten Limits, der Warnung oder der Kosten pro Anfrage berücksichtigt wird.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Summary
Provider API, model gpt-5.6-sol, plan individual-goat:
$15.5624 charged across 12 requests in 2 minutes 49 seconds (2026-09-19 05:36:24-05:39:13 local UTC+8 / 2026-09-18 21:36-21:39 UTC).
Each request carried ~305k-336k tokens of context, giving an effective rate of ~$4.02 per 1M tokens. This single 3-minute burst exceeded the 5-hour window cap ($14; window shows $17.2783 used) and consumed 44% of the $35 weekly cap.
Expected Behavior
- A request that returns FAILED should not be billed in full.
- Exceeding a 5-hour window should not be reachable by one client in under 3 minutes with no warning.
- A per-API-key spend cap or alert (and/or a warning threshold before the window is exhausted) would prevent silent burn.
Actual Behavior
- The request at 05:39:13 returned FAILED and was still billed $0.2844.
- A background agent made 10 calls on gpt-5.6-sol costing ~$7.19 with no usable output at all.
- 4 of the 12 requests cost ~$3.00-3.31 each ($12.6744 = 81.4% of the spend); the other 8 cost $0.27-0.44. The expensive ones appear to be full cache-prefix rewrites billed at $6.25/1M.
- Net: $15.5624 for 3,869,607 tokens (1,306,770 input + 2,562,837 cache read + 14,367 output) = $4.02 per 1M tokens.
Steps to reproduce the issue
- Use the Provider API (api.commandcode.ai/provider/v1) with model gpt-5.6-sol on a plan that has a 5-hour window cap.
- Send requests from a client that carries a large context (~300k+ tokens per call) and does not keep the prompt prefix stable - e.g. a harness that compacts context, retries, or switches models.
- Observe ~$1.30 average per request, then the 5-hour window being exceeded within minutes.
Command Code Version
1.56.2
Operating System
Windows
Terminal/IDE
DeepSeek Harness (dsh) via Provider API, not the CLI
Shell
PowerShell
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
Requesting a review and credit for the FAILED request ($0.2844) and the 10 no-op background-agent calls (~$7.19) - failed or empty runs should not bill at full rate.
Account context:
- plan: individual-goat; billing period started 2026-09-18T12:02:57Z
- official usage summary for the period: $19.2692 over 1,573 requests, 670,391,454 input tokens / 1,170,128 output tokens
- /alpha/billing/credits at the time: fiveHour used $17.2783 / cap $14 (exceeded), weekly used $19.2547 / cap $35
Feature request that would have prevented this: expose per-request cost in the API (or a spend webhook) so clients can stop before burning a whole window. Note also that the usage API does not break out cache-write tokens, which makes client-side cost meters under-report roughly 2x for models that bill cache writes.
Happy to share the 12 request IDs from the Studio usage log.
- Vorherrschende Sprache
- Keine Sprachdaten
- Sterne
- 4k
- Forks
- 350
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus CommandCodeAI/command-code
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#855 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
CommandCodeAI/command-code#841 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#655 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#608 ·
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 70/100
CommandCodeAI/command-code#893 ·
Alle Issues in CommandCodeAI/command-code
Ähnliche Issues
-
enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
JuliusBrussee/caveman#1102 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
use-agent-os/agent-os#3263 ·
-
[Bug]: context-limit error parsing has no pattern for llama.cpp's "context size (N tokens)" phrasing Offenarea/compression area/local-models area/sessions comp/agent duplicate P2 sweeper:risk-session-state type/bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
NousResearch/hermes-agent#117793 · 1 Kommentar ·
-
possible bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Mintplex-Labs/anything-llm#6415 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100