openai / openai/codex

Completed replies missing while Codex Windows task remains on Thinking

Open
#45,702 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

app app-server bug windows-os
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of the Codex App are you using (From “About Codex” dialog)?

26.908.4834.0

What subscription do you have?

Not Sure - I think Professional?

What platform is your computer?

Microsoft Windows NT 10.0.26200.0 x64

What issue are you seeing?

Completed replies missing while Codex Windows task remains on Thinking

Summary

Codex on Windows intermittently keeps a task on Thinking, with the stop button and sidebar activity indicator visible, after the assistant has completed and saved its reply. A completion notification appears, but the reply does not appear in the conversation. The user reports this happens frequently and recovery requires closing all Codex windows and fully exiting remaining ChatGPT apps from the system tray before reopening.

Status: Evidence collected on 2026-09-15; not submitted externally. Completion and persisted final answers confirmed. User subsequently tested Open in New Window: the new desktop view is also stuck. Android screenshot now confirms the final reply is visible in the same named task with remote host mrybicki-hp shown. Full app exit/reopen consistently restores the desktop view according to the user. Root cause remains unverified.

Impact: Repeated interruptions, uncertainty about whether work finished, and full-app restart needed to read completed answers. Frequency is user-reported as frequent; no occurrence count or exact first affected release established.

Environment

  • Device: MRYBICKI-HP.

  • Windows 11 Enterprise, 10.0.26200 / build 26200.

  • Installed Codex package: OpenAI.Codex 26.908.4834.0.

  • Desktop logs report release 26.908.40834. Package and application release identifiers are distinct; About-screen version not yet captured.

  • ChatGPT Classic is also running; package OpenAI.ChatGPT-Desktop 1.2026.190.0.

  • Codex package processes are named ChatGPT.exe. Classic uses ChatGPT Classic.exe. Multiple processes alone do not prove multiple independent app instances or a defect.

  • Time zone: America/New_York, UTC-04:00 on the incident date (EDT). Windows names the zone Eastern Standard Time even when daylight saving is active.

  • Target task: _MIRA 2026-09-11+.

  • Task link: codex://threads/01a091c9-38b9-7393-b172-357f764412c4

  • Target ID: 01a091c9-38b9-7393-b172-357f764412c4.

  • Workspace: C:\Github. Screenshot shows GPT-6 Astra Light.

Expected and actual behavior

Expected: When the final reply is persisted and the turn completes, every open view of the task displays the reply, clears Thinking, and restores the composer to its idle state.

Actual: Screenshot shows completed-work duration rows with no visible answers, followed by the 8:59 AM user message, Thinking, and an active stop button. Saved session and task API independently show completion. User saw a notification but could not read the response in the task.

Evidence and timeline

The original screenshot is screenshot-stuck-thinking.png. Its exact capture time is not established; displayed message times must not be treated as screenshot capture time.

The follow-up screenshot, screenshot-sidebar-spinner.png, highlights the spinner next to _MIRA 2026-09-11+ while a different task (this investigation) is selected. The user reports that the highlighted wheel continued spinning indefinitely. A still image confirms its presence, while duration/animation are user-reported. This broadens the symptom to stale sidebar activity across task navigation. The separate spinner on the active investigation is expected during collection and is not evidence of another stuck task. The adjacent clock-like icon is not being classified as a fault.

Time on 2026-09-15 (EDT) | UTC | Evidence -- | -- | -- 07:34:10 | 11:34:10 | Morning scheduled turn has persisted final reply and task_complete. 08:33:33.501 | 12:33:33.501Z | Journal-correction response persisted as assistant message, phase=final_answer. 08:33:33.539 | 12:33:33.539Z | task_complete contains the correction response. 08:33:33.619 | 12:33:33.619Z | Desktop log records show turn-complete for the same task/turn. 08:59:17.025 | 12:59:17.025Z | Follow-up turn starts; user asks whether the assistant is stuck. 08:59:20.350 | 12:59:20.350Z | Renderer window 1 logs ResizeObserver loop completed with undelivered notifications. Correlation only. 08:59:21.781 | 12:59:21.781Z | Final answer persisted: “No—I’ve finished correcting the journal.” 08:59:21.838 | 12:59:21.838Z | task_complete persists the same final reply. 08:59:21.903 | 12:59:21.903Z | Desktop logs show turn-complete notification for the target turn, renderer window 4 / web contents 5, visible but unfocused. 08:59:21.907 / .971 | 12:59:21.907 / .971Z | IPC warnings: Received broadcast but no handler is configured, method=thread-read-state-changed. Approximately 09:25 | Approximately 13:25Z | Read-only task API returns idle and latest 3 turns completed with null errors, but items=[] for all 3 even with includeOutputs=true. 09:28:39 | 13:28:39Z | Successful baseline evidence snapshot; 4 app logs, 366 task-linked log lines, 12 session records; 0 collection errors.

The phone screenshot has been received and preserved. The report now includes visual evidence from both clients and saved completion records. A simultaneous screenshot pair or Android app version can supplement the report but is not required to submit the current evidence. Repeated restart tests are not required to establish the already reported workaround.

Reproduction scenario to validate

  1. Use the Windows desktop app with a long-lived task and multiple app windows. Record which windows show the target task.

  2. Let a scheduled run or typed follow-up finish while changing window focus.

  3. Observe whether a completion notification appears while a task view remains on Thinking and lacks the answer.

  4. Compare all open views and capture persisted completion state before restarting.

  5. Test switching away/back, reopening only the affected window, and finally fully exiting Codex, recording the first action that restores the answer.

These are candidate reproduction conditions based on this incident, not a proven deterministic recipe. A scheduled run is visible in the screenshot; neither scheduling nor voice has been established as required.

Prior history

The existing local report C:\Github\Codex-Windows-Archive-and-Voice-Bug-Report.md, dated 2026-08-22, was reopened during this investigation. It already describes a voice-to-text conversation staying on Thinking until a full restart, on package 26.818.2872.0. Its archive failures are separate historical symptoms. Similarity supports recurrence of the visible symptom; it does not prove an identical root cause or uninterrupted presence in every intervening release.

Evidence package

  • screenshot-android-completed.jpg and android-screenshot-provenance.json: phone view of the completed reply, original-copy hash verified.

  • thread-api-snapshot.json: live API result; unrelated historical task preview omitted.

  • captures/20260915-092839-421-baseline/environment.json: app packages, OS, process paths/parents/start times; no command lines.

  • Same capture, session-completion-evidence.jsonl: selected incident-day user/final/completion records with original source path and line number. Tool output and reasoning excluded.

  • Same capture, recovered-replies.md: the 3 completed answers. Baseline heading times are UTC despite missing explicit zone labels; the collector was subsequently corrected to emit ISO UTC headings.

  • Same capture, private-app-logs/: 4 app-log snapshots from September 14–15. Active files can grow while copied; hashes describe copied bytes, not a transactionally consistent global snapshot.

  • Same capture, log-provenance.json: source paths, sizes, snapshot hashes.

  • Same capture, thread-log-excerpts.txt: 366 lines matching the target ID. Warnings without IDs remain in the full log snapshots.

  • diagnostic-excerpts.txt: focused completion/error/IPC lines with source line numbers, excluding account-identifying notification-forward records.

  • Capture-Evidence.ps1: manual repeatable snapshot collector. No continuous/background tracing enabled.

  • Troubleshooting.md: next tests and outcome ledger.

Final collector verification: captures/20260915-093051-157-verified/ contains 4 logs, 366 task-linked lines and 12 session records with 0 collection errors. All 4 saved-log SHA256 hashes were read back and matched. Recovered-reply headings in this capture explicitly use ISO UTC timestamps. Both original screenshots were copied without editing.

Collection attempts named initial and initial-retry were incomplete during collector development (shared-file read failure, followed by date-filter/path bugs). Do not use them as evidence of absent logs or replies. The baseline capture is the first validated capture. Existing attempts were preserved, not deleted.

Sharing and support request

Please investigate why persisted completed turns and desktop notifications do not update the task view, focusing on renderer ownership/follower subscriptions, item retrieval and IPC invalidation. Correlate the target IDs and UTC timestamps above with internal telemetry.

The screenshot, recovered replies, and private logs include business context, paths and identifiers. Keep the bundle local; review/redact attachments before external sharing. No authentication/config files or database copies were collected. No report has been posted and no session sharing has been enabled.

Official documentation describes feedback from the / composer menu, optional session sharing and a returned session ID. It does not establish a fix for this particular bug: https://learn.chatgpt.com/docs/reference/troubleshooting#feedback-and-logs (accessed 2026-09-15).

Codex-DisplayRefresh-Full-Evidence.zip

What steps can reproduce the bug?

Reproduction scenario to validate

Use the Windows desktop app with a long-lived task and multiple app windows. Record which windows show the target task.

Let a scheduled run or typed follow-up finish while changing window focus.

Observe whether a completion notification appears while a task view remains on Thinking and lacks the answer.

Compare all open views and capture persisted completion state before restarting.

Test switching away/back, reopening only the affected window, and finally fully exiting Codex, recording the first action that restores the answer.

These are candidate reproduction conditions based on this incident, not a proven deterministic recipe. A scheduled run is visible in the screenshot; neither scheduling nor voice has been established as required.

What is the expected behavior?

Expected and actual behavior

Expected: When the final reply is persisted and the turn completes, every open view of the task displays the reply, clears Thinking, and restores the composer to its idle state.

Actual: Screenshot shows completed-work duration rows with no visible answers, followed by the 8:59 AM user message, Thinking, and an active stop button. Saved session and task API independently show completion. User saw a notification but could not read the response in the task.

Additional information

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with thread-log-excerpts.txt, diagnostic-excerpts.txt, and thread-api-snapshot.json to compare persisted completion, notifications, and the stale desktop views for the target ID. Investigate the renderer ownership or follower subscriptions, item retrieval, and IPC invalidation mentioned in the report; done means a completed reply appears, Thinking clears, and the composer returns idle in every open view without a full app restart.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
desktop, devtools
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.