Chat responses are consistently much shallower than ChatGPT Web with the same prompt and reasoning settings
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.5k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
Chatgpt Pro
What subscription do you have?
Chatgpt Pro 20x
What platform is your computer?
No response
What issue are you seeing?
Additional information
This report is specifically about a systematic Desktop vs Web discrepancy, not normal stochastic variation between model responses.
I have already tested the same prompts across both clients, and the difference is consistently large enough to affect practical usability.
I am also aware that response latency alone is not sufficient evidence of a model downgrade. My concern is based on the combination of:
repeated identical-prompt comparisons,
materially reduced analysis depth,
consistently poorer task completion in Desktop,
and the fact that similar reasoning/configuration propagation bugs have previously been documented in Codex Desktop.
For that reason, I believe the most useful investigation would be to compare the server-side telemetry of an affected Desktop turn with the equivalent Web turn, especially the requested model, served model, requested reasoning effort, effective reasoning configuration, and any routing/fallback metadata.
What steps can reproduce the bug?
Feedback ID: no-active-thread-019ff45d-7f3d-7563-b66c-e30eab9599cd
What is the expected behavior?
No response
Additional information
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with feedback ID no-active-thread-019ff45d-7f3d-7563-b66c-e30eab9599cd and compare the affected Desktop turn with the equivalent Web turn. Investigate requested and served models, requested and effective reasoning settings, and routing or fallback metadata in server-side telemetry. Done means identifying the source of the consistent response-depth discrepancy.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- desktop-dev, observability
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100