openai / openai/codex

Chat responses are consistently much shallower than ChatGPT Web with the same prompt and reasoning settings

Open
#38,125 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

app bug model-behavior
Dominant language
Rust
Stars
125k
Forks
19.5k
PR merge metrics
PR metrics pending

Description

What version of the Codex App are you using (From “About Codex” dialog)?

Chatgpt Pro

What subscription do you have?

Chatgpt Pro 20x

What platform is your computer?

No response

What issue are you seeing?

Additional information

This report is specifically about a systematic Desktop vs Web discrepancy, not normal stochastic variation between model responses.

I have already tested the same prompts across both clients, and the difference is consistently large enough to affect practical usability.

I am also aware that response latency alone is not sufficient evidence of a model downgrade. My concern is based on the combination of:

repeated identical-prompt comparisons,
materially reduced analysis depth,
consistently poorer task completion in Desktop,
and the fact that similar reasoning/configuration propagation bugs have previously been documented in Codex Desktop.

For that reason, I believe the most useful investigation would be to compare the server-side telemetry of an affected Desktop turn with the equivalent Web turn, especially the requested model, served model, requested reasoning effort, effective reasoning configuration, and any routing/fallback metadata.

What steps can reproduce the bug?

Feedback ID: no-active-thread-019ff45d-7f3d-7563-b66c-e30eab9599cd

What is the expected behavior?

No response

Additional information

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with feedback ID no-active-thread-019ff45d-7f3d-7563-b66c-e30eab9599cd and compare the affected Desktop turn with the equivalent Web turn. Investigate requested and served models, requested and effective reasoning settings, and routing or fallback metadata in server-side telemetry. Done means identifying the source of the consistent response-depth discrepancy.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
desktop-dev, observability
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.