[Windows][Plus] Astra low consumes 90% of five-hour quota in ~10 minutes; 99.68% cached input
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.908.4834.0 (installed Windows package); bundled codex-cli 0.154.0-alpha.6.2
What subscription do you have?
ChatGPT Plus
What platform is your computer?
Windows 11 Pro, 10.0.26200, x64
What issue are you seeing?
Recently my five-hour Codex allowance started depleting in roughly 10 minutes, including when selecting low reasoning. I observe this across projects and conversations. Please investigate whether this is an accounting regression or unexpectedly high context charging.
Verified local evidence on September 12, 2026 (UTC):
- Existing Astra low session: 14:22:42.303, five-hour used=5%; 14:33:00.000, used=95%. That is 90 percentage points in approximately 10m18s.
- Cumulative-token differences across those endpoints: input 8,405,735; cached input 8,378,496 (~99.68% of input); output 7,532; reasoning output 944 (included in output, not added separately).
- Another low-reasoning interval that morning rose from 5% at 08:47:57.994 to 99% at 09:01:32.590.
- A brand-new project/conversation created specifically to investigate confirms model=gpt-6-astra, effort=low. First inference had 27,975 input tokens, 19,456 cached input, 173 output and zero reasoning output before substantive investigation. Its quota has increased from approximately 0–1% to 9% during the early diagnostic steps; full exhaustion has NOT yet been reproduced in this fresh conversation.
- Current config has service_tier="default" and model_reasoning_effort="low". Global AGENTS.md is empty. The fresh diagnostic turn has spawned no subagents.
These are account-wide server-reported quota snapshots recorded alongside local token events, not a per-task billing ledger. I cannot establish that all account usage belongs to this one task. Cached-context reuse itself is expected and these counts do not prove duplicate billing. The older session has substantial context; the new conversation still has app/tool/skill overhead. One active heartbeat exists locally, and its contribution has not been isolated. No reset credit was redeemed during this investigation. No local fix has been established. No raw logs, credentials, account IDs or project contents are included.
What steps can reproduce the bug?
- Use Codex desktop on Windows with a Plus subscription.
- Select Astra and low reasoning with default service tier.
- Run a task involving repeated tool calls and observe the five-hour allowance.
- Compare token_count events and turn_context model/effort with server quota snapshots.
- Also start a fresh project and conversation to distinguish accumulated conversation context from baseline app/tool overhead.
Historical measured intervals above reproduce rapid depletion; the fresh diagnostic conversation is a partial observation, not yet a reproduction of 100% depletion. Exact onset and a controlled minimal reproduction remain unconfirmed.
What is the expected behavior?
Usage should be consistently accounted for, with documented model/reasoning/cache treatment and enough visibility to explain a sudden change. Please audit the quota ledger, cached-input weighting, possible duplicated/auxiliary charges and entitlement/window reconciliation. If incorrect charges are confirmed, restore the affected allowance. I am not claiming that a five-hour window guarantees five hours of continuous generation.
Additional information
Related reports: #45067 (cross-model depletion), #44386 (cached context and Plus usage), #45073 (rapid five-hour depletion). This report adds measured Windows/Astra-low timing and cache totals. Please consolidate if these share a root cause.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by comparing the reported token_count events and turn_context model/effort values with the server quota snapshots, using the existing Astra-low intervals and fresh conversation as baselines. A useful resolution would establish a controlled minimal reproduction and distinguish accumulated context, cached-input treatment, app/tool overhead, heartbeat activity, and account-level quota reconciliation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- devtools
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100