[macOS Desktop] Work/Codex tasks give unclear product-mode explanations and fail to distinguish native behavior from customization
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.908.40834 (build 8881).
What subscription do you have?
Not included; no subscription-specific failure is alleged.
What platform is your computer?
macOS 26.6.2, arm64. Unified ChatGPT/Codex desktop app.
What issue are you seeing?
The assistant in a desktop task failed to give a clear, reliable explanation of the difference between opening a task under ChatGPT Work and Codex, including whether the differences were native product behavior or user customizations.
Observed September 13, 2026:
- Asked for an exact comparison, the assistant initially gave a generic capability comparison and told the user to check project, files, instructions, tools and model inheritance.
- The user had to demand current official documentation and repository research to obtain a concrete explanation.
- The later answer introduced
STEPS_PROSEandSTEPS_COMMANDSwithout plainly identifying these as native app implementation labels rather than user-written customizations. The user again had to ask for that distinction.
The failure is in the assistant's product grounding and explanation, not merely the visible menu labels. Users should not have to investigate app internals or repeatedly correct the assistant to understand the product surface in which they are working.
This is an observed answer-quality failure in one interaction. It is not a controlled clean-profile A/B test, does not establish that every Work or Codex task fails, and does not establish a backend root cause.
What steps can reproduce the bug?
Observation sequence and a suggested evaluation prompt:
- Open the desktop app and inspect the ChatGPT/Codex selector.
- In a desktop task, ask: “Exactly what changes when I open a task under ChatGPT Work versus Codex in this app?”
- Ask: “Which differences are native app behavior or built-in instructions, and which come from my personal or project customizations?”
- Check whether the initial answer separates ordinary Chat, local Work, cloud Work and Codex, and attributes each claimed difference to the correct layer.
In the observed interaction, multiple corrections were needed before the native-versus-custom distinction was explained. Reproducibility across fresh tasks, accounts and builds remains untested.
What is the expected behavior?
The first answer should:
- Distinguish ordinary Chat from Work, local desktop Work from cloud Work, and local Work from Codex.
- Explain which differences concern native UI presentation and built-in behavior instructions, which concern execution/environment, and which concern optional personal/project customization.
- Treat model, reasoning effort, permissions, local file access, history/memory and usage as separate dimensions.
- Use current official documentation and available read-only task metadata.
- Name a specific unavailable fact when necessary, while still explaining the documented product distinction.
- Avoid treating a generic Codex agent identity as sufficient evidence of the selected app mode.
Requested improvement: strengthen model-visible product/mode grounding and official-documentation retrieval, expose structured current mode and execution/capability metadata through supported read-only controls, and add evaluations for accurate Work/Codex comparisons and native-versus-custom attribution.
Additional information
Official reference:
https://learn.chatgpt.com/docs/use-chatgpt#compare-work-mode-and-codex-on-desktop
Related but distinct reports:
- #33079 requests persistent Work/Codex badges and conversation identification.
- #41594 concerns generic “chat” labels in creation controls and navigation.
Those UI changes alone would not fix an assistant that cannot accurately explain native versus customized behavior.
Feedback receipt: no-active-thread-01a09bbe-48ac-7d43-9391-cfc1444c7477
Acceptance: On an affected desktop build, fresh tasks in both local Work and Codex should accurately explain the documented comparison and correctly distinguish native behavior from optional customization without repeated user corrections. Any unavailable current-task metadata should be identified narrowly. Backend cause and cross-account frequency remain NOT_PROVED.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the official Work/Codex comparison linked in the issue, then reproduce the suggested prompts in fresh local Work and Codex tasks. Inspect supported read-only task metadata and the product-grounding or evaluation entry points that are available in the repository. Done means the assistant separates modes, execution context, native behavior, and optional customization, while naming unavailable facts narrowly without repeated correction.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- macos, rust
- Domain
- ai, desktop, testing-qa
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 42/100