Astra reasoning quality appears significantly degraded compared with previous sessions
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.908.40834
What subscription do you have?
ChatGPT ProX20
What platform is your computer?
Windows 10
What issue are you seeing?
Astra's reasoning quality appears to have degraded significantly. In recent sessions, it performs noticeably worse on the same types of tasks that it previously handled correctly.
The model is more likely to misunderstand explicit instructions, produce shallow or incorrect reasoning, lose important context, and require repeated corrections. The difference is large enough that it feels like the model's effective reasoning capability has been reduced.
This is not just a single bad response. I have observed the behavior repeatedly in actual Codex tasks. I submitted the affected conversation through /feedback so the team can inspect the original thread and full context.
Feedback ID: 01a096e9-26b6-7371-8d76-35dfc8739455
Please check whether there have been changes to Astra's routing, reasoning configuration, context handling, or model behavior that could explain this regression.
What steps can reproduce the bug?
Feedback ID: 01a096e9-26b6-7371-8d76-35dfc8739455
What is the expected behavior?
Fix
Additional information
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the submitted conversation associated with feedback ID 01a096e9-26b6-7371-8d76-35dfc8739455, using the reported Codex App version 26.908.40834 and Windows 10 environment. Compare the affected behavior with prior sessions and investigate whether routing, reasoning configuration, context handling, or model behavior changed. Done means identifying and addressing the reported reasoning-quality regression.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100