continuedev / continuedev/continue

OpenAI cost calculation ignores cached input tokens (over-costs cache hits)

Open Beginner friendly
#13,104 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
TypeScript
Stars
36k
Forks
5.4k
PR merge metrics
No merged PRs in 30d

Description

Summary

calculateOpenAICost in core/llm/utils/calculateRequestCost.ts bills the full promptTokens at the standard input rate and never accounts for cached input tokens, so requests that hit OpenAI's prompt cache are over-costed. The Anthropic branch in the same file already handles cache tokens; the OpenAI branch does not.

Detail

OpenAI reports cached input as usage.prompt_tokens_details.cached_tokens, and prompt_tokens includes those cached tokens. Cached input is billed at a discount (for example gpt-4o cached input is half the standard input rate). The cost function charges every prompt token at full rate:

const inputCost = (usage.promptTokens / 1_000_000) * modelPricing.input;
// no use of usage.promptTokensDetails.cachedTokens

Compare calculateAnthropicCost, which reads usage.promptTokensDetails and prices cachedTokens / cacheWriteTokens at their own rates.

Effect

For an OpenAI request with cached input (common with long, stable system prompts), the reported cost is higher than the actual OpenAI charge — the cached portion is billed at full price instead of the cache-read discount.

Suggested direction

Give the OpenAI pricing table a cachedInput rate and subtract the cached tokens from the full-rate input, pricing them separately, the way the Anthropic branch does:

const cachedTokens = usage.promptTokensDetails?.cachedTokens ?? 0;
const uncachedInput = Math.max(0, usage.promptTokens - cachedTokens);
const inputCost = (uncachedInput / 1_000_000) * modelPricing.input
                + (cachedTokens / 1_000_000) * modelPricing.cachedInput;

Happy to open a PR with the cached-input rates for the models already listed if that direction sounds right.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in core/llm/utils/calculateRequestCost.ts at calculateOpenAICost and compare it with the Anthropic branch's cache-token handling. Review the OpenAI pricing entries and promptTokensDetails usage fields, then verify that cached and uncached input are priced separately and the reported total matches the provider's charge.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, backend
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
74/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.