Codex App: premature completion claims and lost task intent in a long session; please investigate gpt-6-astra / ultra execution
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.901.2854.0
What subscription do you have?
chat gpt pro 20x
What platform is your computer?
Microsoft Windows NT 10.0.19044.0 x64
What issue are you seeing?
I am reporting serious instruction-following and completion-reporting failures in a long-running Codex App session. I also request an investigation into whether the actual model and reasoning settings matched the requested configuration.
The ongoing task was to deliver a ready-to-run ML training package for a colleague's RTX 5090 Laptop with an existing training environment. Downloading a checkpoint and providing a training plan were intermediate steps, not the requested final deliverable.
The assistant described the training package as ready, then later explicitly admitted that only the checkpoint and original source were available. The new sampling recipe, stage-specific prompts, continuation configuration, and launch scripts had not been integrated. It ended its reply with this status explanation rather than completing the outstanding deliverable.
Its subsequent correction included the following, translated from Chinese:
The complete training checkpoint has been downloaded, but the one-command training package for the new recipe is not finished.
This distinction should have been made before reporting readiness. The misleading status and repeated need to restate the same task disrupted a time-sensitive competition handoff.
Some replies also appeared unusually immediate and shallow compared with my expectations for the selected model and reasoning setting. Local session records inspected by the assistant showed model=gpt-6-astra and effort=ultra. The assistant initially presented that as confirmation of its current model, then acknowledged that local configuration alone does not independently establish which model actually served a request.
I cannot verify backend routing. I am asking OpenAI to investigate a possible model/reasoning mismatch, not presenting response speed alone as proof of a fallback.
Additional concrete failure while preparing this very bug report
The assistant initially left the Codex App version as a placeholder and told me to obtain it manually, despite having shell access to the computer where the app was installed. Its first attempt to run Get-AppxPackage failed because the Appx module could not load in the current PowerShell host. It then handed off the report without trying the available Windows PowerShell host.
Only after I challenged this did it query the same machine through Windows PowerShell and retrieve:
Name: OpenAI.Codex
Version: 26.901.2854.0
This was not information that only I could supply, nor was it a CLI-versus-app ambiguity that could not be resolved. It was an avoidable failure to recover from a tool error and complete a basic local lookup. It happened during preparation of a report about the assistant's reasoning and execution quality, further undermining my confidence.
My routing concern is therefore based on several concrete failures, not merely a short response time: misleading completion claims, failure to carry an established deliverable through to completion, and abandoning a straightforward local lookup after the first tool error. From a user's perspective, this feels inconsistent with the selected gpt-6-astra / ultra configuration. Please check whether the affected requests were actually routed to the requested model with the requested reasoning setting, rather than treating the local configuration string as sufficient evidence that they were.
What steps can reproduce the bug?
Feedback ID: 01a031af-a426-7190-b1a2-9a68f1d8e08e
What is the expected behavior?
Do not degraded my model please...
Additional information
- Feedback ID:
01a031af-a426-7190-b1a2-9a68f1d8e08e - Incident date: September 19, 2026; timezone: Asia/Shanghai (UTC+8).
- Requested model / locally recorded effort:
gpt-6-astra/ultra. - Account usage reported at the time: 73% of the weekly allowance used, 27% remaining. This is account-level usage, not this session's context-window utilization.
- Context-window utilization and per-response reasoning-token counts: not verified.
Please investigate:
- The actual model/snapshot serving the affected recent turns, including the bug-report drafting turns, and whether any requests were routed to a different model or affected by a fallback.
- Whether the requested
ultrareasoning effort was accepted and applied, or overridden. Please distinguish request metadata from actual server-side execution when investigating. - Whether context compaction or long-session state handling contributed to loss of the task objective and inaccurate completion claims.
I am not requesting private chain-of-thought. I am requesting verification of execution settings and investigation of the observable failures. If backend verification is outside this repository's scope, please route this report to the appropriate team using the feedback ID.
The original session contains private project information and credentials, so I am not attaching an unredacted transcript to a public GitHub issue.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with Feedback ID 01a031af-a426-7190-b1a2-9a68f1d8e08e and the described local session records, then determine whether this repository exposes any way to verify model routing, reasoning settings, or long-session state. Review the Windows PowerShell lookup details as a separate reproducible symptom. Done means establishing repository scope and routing the report appropriately if backend evidence is unavailable.
Written by the indexing model from the issue text.
Assessment
- Domain
- developer-experience, tooling
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 18/100