CommandCodeAI / CommandCodeAI/command-code

Provider API: failed requests and no-op agent runs are billed in full ($15.56 in under 3 minutes on gpt-5.6-sol)

未關閉
#888 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

主要語言
沒有語言資料
星號
4k
分支
350
PR 合併指標
30 天內沒有已合併 PR

描述

Summary

Provider API, model gpt-5.6-sol, plan individual-goat:
$15.5624 charged across 12 requests in 2 minutes 49 seconds (2026-09-19 05:36:24-05:39:13 local UTC+8 / 2026-09-18 21:36-21:39 UTC).

Each request carried ~305k-336k tokens of context, giving an effective rate of ~$4.02 per 1M tokens. This single 3-minute burst exceeded the 5-hour window cap ($14; window shows $17.2783 used) and consumed 44% of the $35 weekly cap.

Expected Behavior
  1. A request that returns FAILED should not be billed in full.
  2. Exceeding a 5-hour window should not be reachable by one client in under 3 minutes with no warning.
  3. A per-API-key spend cap or alert (and/or a warning threshold before the window is exhausted) would prevent silent burn.
Actual Behavior
  • The request at 05:39:13 returned FAILED and was still billed $0.2844.
  • A background agent made 10 calls on gpt-5.6-sol costing ~$7.19 with no usable output at all.
  • 4 of the 12 requests cost ~$3.00-3.31 each ($12.6744 = 81.4% of the spend); the other 8 cost $0.27-0.44. The expensive ones appear to be full cache-prefix rewrites billed at $6.25/1M.
  • Net: $15.5624 for 3,869,607 tokens (1,306,770 input + 2,562,837 cache read + 14,367 output) = $4.02 per 1M tokens.
Steps to reproduce the issue
  1. Use the Provider API (api.commandcode.ai/provider/v1) with model gpt-5.6-sol on a plan that has a 5-hour window cap.
  2. Send requests from a client that carries a large context (~300k+ tokens per call) and does not keep the prompt prefix stable - e.g. a harness that compacts context, retries, or switches models.
  3. Observe ~$1.30 average per request, then the 5-hour window being exceeded within minutes.
Command Code Version

1.56.2

Operating System

Windows

Terminal/IDE

DeepSeek Harness (dsh) via Provider API, not the CLI

Shell

PowerShell

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

Requesting a review and credit for the FAILED request ($0.2844) and the 10 no-op background-agent calls (~$7.19) - failed or empty runs should not bill at full rate.

Account context:

  • plan: individual-goat; billing period started 2026-09-18T12:02:57Z
  • official usage summary for the period: $19.2692 over 1,573 requests, 670,391,454 input tokens / 1,170,128 output tokens
  • /alpha/billing/credits at the time: fiveHour used $17.2783 / cap $14 (exceeded), weekly used $19.2547 / cap $35

Feature request that would have prevented this: expose per-request cost in the API (or a spend webhook) so clients can stop before burning a whole window. Note also that the usage API does not break out cache-write tokens, which makes client-side cost meters under-report roughly 2x for models that bill cache writes.

Happy to share the 12 request IDs from the Studio usage log.

貢獻指南

這個儲存庫沒有索引到貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 Provider API 入口 api.commandcode.ai/provider/v1 開始,追蹤失敗請求、空代理執行、快取寫入權杖和支出上限是如何計費的。未指定任何儲存庫檔案或測試;完成後應確定失敗呼叫或 no-op 呼叫是否被正確計費,以及回報的上限、警告或每次請求成本行為是否得到處理。

由索引模型根據 Issue 內容生成。

評估

領域
ai, api, payments
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
活躍
描述清晰度
基本清楚
新手友好度
42/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。