Silently switched models and cost extra time and tokens to fix.
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
Research direction
Start with the uploaded thread 019f66dc-8ca7-74e1-a825-d91c7fc7c427 and the model-setting events described in the report; reproduce on Codex CLI 0.144.4 using the selected Luna model. Trace where the forced switch occurs and verify that a user's model is not changed silently during a stream. Done means the selected model remains in use unless the user explicitly changes it.
Written by the indexing model from the issue text.
Description
What version of Codex CLI is running?
codex-cli 0.144.4
What subscription do you have?
$20/month
Which model were you using?
luna light
What platform is your computer?
Darwin 25.5.0 arm64 arm
What terminal emulator and version are you using (if applicable)?
wezterm
Codex doctor report
codex doctor returned overallStatus: "fail".
1. Failure: ChatGPT HTTP endpoint unreachable.
2. Warning: WebSocket DNS lookup failed.
3. Warning: Homebrew update check could not resolve formulae.brew.sh.
4. All auth, config, installation, databases, rollout parity, Git, MCP, and
search checks passed.
5. Installed version: 0.144.4.
Important: this command ran inside Codex’s restricted-network sandbox, so the
network failures likely describe the diagnostic process’s sandbox access—not
necessarily your Mac’s normal connectivity.
What issue are you seeing?
name: reference_oai_model_switch_cost
description: "Verified token-cost ledger for the 2026-07-15 forced Luna-low model switch and remediation"
metadata:
node_type: memory
type: reference
OAI forced-model-switch cost ledger
Source of truth: local Codex rollout JSONL token-count and thread-settings events. Aggregate tokens
include cached input. Reasoning tokens are reported separately by Codex metadata and may be a subset
of output accounting. Do not represent aggregate tokens as an API bill without OpenAI's pricing and
cache treatment; they are the exact product-reported consumption counter.
Incident boundary
- Pre-switch counter, 2026-07-15 15:25:25 EDT: 15,195,167 total tokens; weekly limit 11%.
- Forced setting applied, 15:25:46 EDT:
gpt-5.6-luna, reasoning effortlow. - Restored setting applied, 16:01:30 EDT:
gpt-5.6-sol, reasoning effortmedium. - First restored-model counter, 16:01:30 EDT: 27,401,495 total tokens; weekly limit 23%.
Forced-Luna interval consumption: 12,206,328 aggregate tokens:
- Cached input: 11,630,080.
- Uncached input: 563,303.
- Output: 12,945.
- Reasoning output reported separately: 4,130.
Remediation snapshots
- 16:04:05 EDT: 28,613,671 total; remediation since restoration 1,212,176; incident total
13,418,504; weekly limit 25%. - 16:06:56 EDT: 29,261,044 total; remediation since restoration 1,859,549; incident total
14,065,877 aggregate tokens; weekly limit 27%. - 16:08:46 EDT, immediately before the full repair pass: 31,016,552 total; weekly limit 28%.
- 16:15:51 EDT, after tracker reconstruction, corrective logging, commit/push, and Sisyphus
reconciliation: 31,973,633 total; repair-pass cost 957,081; remediation since restoration
4,572,138; full incident cost 16,778,466 aggregate tokens; weekly limit 32%.
What the remediation is correcting
- False ChatGPT 366th-record repair claim.
- Group C triage vs distillation conflation.
- Raw-Markdown screenshot misdiagnosis.
- Obsidian symlink visibility failure and unverified success claim.
- Duplicate visible/hidden tracker sources of truth.
- Incomplete five-row lifecycle matrix followed by an unsynchronized 19-row revision.
- Inconsistent lifecycle-stage semantics and stale three-stage pipeline text.
- Mac/Sisyphus/GitHub divergence and untracked Group C residue.
- Append-only project-log discipline violations.
Update this ledger from the rollout JSONL after each remediation block. Keep forced-interval cost and
remediation cost separate, and report the weekly-limit percentage change.
What steps can reproduce the bug?
Uploaded thread: 019f66dc-8ca7-74e1-a825-d91c7fc7c427
What is the expected behavior?
Don't silently switch models on your users mid stream.
Additional information
I did not choose to switch models.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.5k
- Avg merge
- 1m
- Merged PRs (30d)
- 1k
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from openai/codex
-
enhancement remote
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
bug CLI windows-os
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
macOS sandbox blocks hw.optional.arm64 sysctl, causing Flutter to misdetect Apple Silicon as x64 Openbug CLI sandbox
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
bug CLI TUI
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
CLI config enhancement skills
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
kwakseongjae/auto-hwp#319 ·
-
area:cli bug filter-quality good first issue priority:medium
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
bevyengine/bevy#25861 ·
-
comp-datalake
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
ClickHouse/ClickHouse#121222 ·
-
A-linter
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
oxc-project/oxc#26863 ·