Windows Desktop: Possible multi-release regression affecting app-server stability, tool execution and session continuity
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.826.12353
What subscription do you have?
ChatGPT Plus
What platform is your computer?
Microsoft Windows NT 10.0.26200.0 x64
What issue are you seeing?
Codex Desktop was previously stable for sustained professional development work on the same Windows environment. More recently, repeated unexpected interruptions have occurred during active, tool-intensive development sessions.
The concern is not an isolated crash, but a possible multi-release Windows Desktop reliability regression.
Observed locally:
- repeated unexpected Codex interruptions during development work;
- interrupted active sessions requiring recovery;
- loss of reliable development continuity after failures;
- session/continuation problems, including 404-related failures and problems occurring around long-running/compacted sessions;
- recurrence across normal, sustained Codex development workflows.
The apparent pattern is:
previously stable workflow → newer release family → repeated failures → recovery required → similar independent Windows reports → problems continuing across subsequent Desktop releases
Several public Codex issues describe overlapping Windows symptoms, including app-server termination/restart during tool execution, missing tool-call outputs, session/thread recovery failures, HTTP 404 behavior, compaction anomalies and abnormal session-state/resource growth.
I am not claiming that these symptoms share one root cause. Some may be independent defects. The concern is whether they are being correlated internally as a broader Windows Desktop reliability regression.
What steps can reproduce the bug?
There is currently no single deterministic minimal reproducer.
The failures occur intermittently during sustained professional Codex development sessions involving local repositories and repeated tool/shell operations.
Typical workflow:
- Open an existing software-development project in Codex Desktop.
- Continue a development session involving repository inspection, file modifications, shell/PowerShell commands and other tool calls.
- Work normally for an extended period, including multiple agent/tool turns.
- During an active development session, Codex may unexpectedly stop or lose the active session/turn.
- Recovery or restart is then required before development can continue.
- In some cases, session/continuation problems, including 404-related failures, have also occurred during long-running development sessions.
The failure is intermittent rather than tied to one specific command.
This is why I suspect a regression/reliability issue rather than a deterministic failure of a particular tool.
A useful diagnostic approach may be to correlate Windows telemetry across:
Last Known Good build → First Known Bad build → active tool execution → app-server lifecycle → session persistence → recovery/compaction
I can provide additional local diagnostic information or logs if Engineering identifies specific data that would help isolate the regression.
What is the expected behavior?
Codex Desktop should remain stable during sustained professional development sessions and preserve transactional development continuity if an internal component fails.
In particular:
- active tool calls should either complete and remain recorded or be clearly marked as interrupted;
- completed tool outputs should not be lost after recovery;
- repository changes should remain deterministically identifiable;
- recovery should not silently duplicate completed operations;
- sessions/threads should remain resolvable after an interruption;
- context compaction should not compromise session integrity;
- an app-server restart should not corrupt or invalidate the active development state.
For professional agentic development, the expected reliability property is not merely that the Desktop window remains open, but that development can continue from a deterministic and recoverable state.
Additional information
Environment
- Windows 11 25H2
- Codex Desktop
- ChatGPT Plus
- Professional sustained software-development workflow
- Local repositories with repeated shell/tool execution
Why I am reporting this as a possible regression
Codex had previously been stable on the same general development environment.
The significant observation is therefore not simply that Codex can crash, but:
previously stable environment + later release family + new recurring failure pattern + similar independent Windows reports + persistence across subsequent builds
Public reports that appear relevant include:
-
#39964 – Windows Desktop failure during tool execution
-
#40231 – different stability characteristics across successive Desktop builds / possible regression window
-
#40323 – automatic compaction associated with extreme rollout growth
-
#40400 – repeated app-server restart and missing tool outputs
-
#40607 – active tool-turn failure associated with transport retries/app-server replacement
-
#41236 – app-server restart with missing custom tool-call outputs
-
#41268 – Windows crash behavior reported as persisting across newer Desktop releases
-
#41333 – task/session history corruption after a Windows interruption, with failed recovery and thread-store inconsistency
These references are provided for correlation only.
I am not asserting that they have a common root cause.
Local observation vs. independent reports
| Observation | Seen locally | Independently reported | Root cause established |
|---|---|---|---|
| Unexpected Codex interruption | Yes | Yes | No |
| Repeated development-session failures | Yes | Yes | No |
| Session recovery required | Yes | Yes | No |
| 404 / compaction-related problems | Yes | Yes | No |
| app-server process replacement | Not locally instrumented | Yes | No |
| Extreme rollout growth (>16 GiB) | No | Yes | No |
| Common root cause | — | — | No |
Questions for Codex Engineering
-
Are these Windows Desktop reports being correlated internally as a broader reliability regression?
-
Has a Last Known Good / First Known Bad build range been identified?
-
Is there currently a recommended Known Good Windows Desktop build for sustained development work?
-
Has any causal relationship been established between app-server restart, tool-output loss, session recovery/404 failures and compaction/session-state behavior?
-
Is there a supported way for affected Windows users to remain temporarily on a qualified stable build while this is investigated?
A Windows-specific Known Issues notice identifying affected builds, known-good build, workaround, fix status and fixed-in version would be very useful for professional users.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No source file, test, or deterministic reproducer is named. Start by comparing the Last Known Good and First Known Bad Desktop builds and reviewing the referenced Windows issues for overlap in app-server lifecycle, tool execution, session persistence, recovery, and compaction. Done means establishing whether a regression window or known-good build can be identified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- powershell
- Domain
- desktop-dev, devtools, operating-systems
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100