openai / openai/codex

[Windows App][Astra] Latest user request visible in UI is missing from rollout after interrupted turns and compaction

Open
#42,930 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

app bug context model-behavior session windows-os
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of the Codex App are you using?

Windows package OpenAI.Codex version 26.901.5003.0, read from Get-AppxPackage without opening a window. The affected Desktop rollout records embedded CLI version 0.153.1. The separately installed CLI currently reports 0.153.4; these are different version observations.

What subscription do you have?

Pro. The shared weekly usage window was at 19% used around the incident, with no rate-limit error identified in the selected event sequence.

What platform is your computer?

Windows x64, Microsoft Windows NT 10.0.26200.0, PowerShell, timezone America/Sao_Paulo (UTC-3).

What issue are you seeing?

A newer user request is visible in the Desktop UI screenshot but absent from the persisted pre-compaction user messages and the replacement history. After two interrupted turns and a compaction, Astra resumes an older task. The user had explicitly asked Codex to gather recent Astra rollouts and Hermes Titan conversation evidence and file this regression report. Codex instead resumed changing a launcher from the previous request. The user had to provide a screenshot of the missing instruction to restore the intended task.

This is also a report of degraded instruction following in long-running Astra work: unrequested restrictions, treating agent-authored history as fact, and repeatedly explaining instead of completing authorized work. The strongest independently inspectable symptom here is the UI/transcript disagreement. The evidence does not establish whether its cause is Desktop submission/persistence, interruption handling, compaction, or model behavior.

Exact incident timeline

Affected session: 01a06e8d-f9b0-78a2-ba73-36caca41c769. The relevant turn context records gpt-6-astra, effort high, provider openai. Earlier turns in this session used medium effort. Times below are UTC on 2026-09-05; line numbers refer to the local JSONL snapshot.

UTC Evidence Observation
04:48:22.714 line 4401, task_complete Agent reports a launcher repair that suppresses commands/output and reasoning summaries from live status. The user had not requested that suppression.
04:52:12.860 line 4405, user message User asks why visibility was removed and when they had asked for that.
04:52:43.878 line 4413, task_complete Agent acknowledges inventing the restriction.
After that answer User-provided Desktop screenshot A newer request asks to gather recent Astra rollouts and Titan history and open an OpenAI issue about the regression. This request is absent from the pre-compaction response_item user messages inspected locally. Its exact submission timestamp is not available in those messages.
04:54:51.365 to 04:57:45.474 lines 4415, 4417 Turn 01a06feb-5ea1-7c20-9b38-51fed1ae72ed starts, then turn_aborted, reason interrupted, duration 174109 ms.
04:57:47.451 to 04:57:51.473 lines 4419, 4421 Turn 01a06fee-0e8b-77f2-9f85-037a3356a183 starts, then turn_aborted, reason interrupted, duration 4022 ms.
04:57:58.171 line 4423 A third turn starts.
05:00:44.131 line 4425, compacted replacement_history retains 23 user messages. The last is the older question about hidden commands. The request to file an issue is missing.
05:00:46.989 line 4435 User says "Travou" ("It froze").
05:00:52.483 line 4440 Agent announces it will finish the previous launcher adjustment, rather than the issue-report task.
05:01:01.323 line 4446 User reports compaction stalls too. Agent continues the launcher task.
05:02:54.314 line 4473 User supplies the screenshot showing the ignored request.
05:03:02.437 line 4478 Agent finally identifies the omitted issue-report request and resumes it.

The screenshot displays multiple automatic-compaction notices. The JSONL contains five compacted records across the whole session; only one is recorded in the narrow interruption window above. The UI notices must not be counted as independently verified successful compactions. reason=interrupted does not identify who or what initiated interruption, so this report does not label those events as backend timeouts.

Before this window, the last recorded inference used 217,209 input tokens against a reported 258,400-token model context window. These are per-call values, not the cumulative session billing total. After compaction, the next recorded inference used 38,592 input tokens.

Additional recent Astra evidence

Four recent local Astra rollouts were inspected, without launching new model sessions:

  • The affected Desktop session above contains both the unsolicited visibility restriction and the omitted latest request.
  • Desktop session 01a06dda-54ce-7e61-8cb5-d02ac41db5fd, Astra medium, recorded CLI 0.153.1: at 01:07:52 UTC the agent presented an "emergency restoration" as historical fact. The user challenged it at 01:12:35. At 01:13:03 the agent acknowledged copying a label from an agent-written handoff and reported that the original user request was a parallel installation. This is a separate source/authority attribution error, not evidence of the same persistence bug.
  • CLI session 01a06f37-bb05-7d33-bc6c-0bb15d4fd2d3, Astra high, CLI 0.153.4: production activation remained incomplete. The executor identified contradictory delegation instructions and a genuinely missing release gate. This is an orchestration confounder; obeying that explicit prohibition is not by itself a model failure.
  • Desktop review session 01a06f6b-f7a0-7f82-8b59-dab34840eda0, Astra medium, recorded CLI 0.153.3: review detected a real bypass, and a later response reported 38 passing tests with a limited local-validation claim. It is included as a counterexample, not counted as another failure.
Hermes Titan evidence, separate third-party environment

The requested review also covered a Hermes fork's Telegram conversation on September 4 and September 5 through 00:31 BRT. It uses openai-codex; recorded main-model usage transitions from gpt-5.6-sol to gpt-6-astra during September 4. It includes custom workflow enforcement, ai-memory, Graphify, large prompts and its own compression implementation. It is not a clean Codex baseline.

The incident audit found:

  • At 22:35:32 BRT, an Astra turn ends with text_response(finish_reason=stop), api_calls=3/200, after an authorized fix and an intervening status/solution question. Answering that question was relevant, but execution resumed only after another user prompt. At 23:33:59 another ordinary stop occurs at 2/200 with delivery still pending.
  • The orchestrator replaced Astra with Sol after a CLI error explicitly said Astra required a newer Codex version. A user correction led to a CLI update and successful Astra probe.
  • A production instruction was delegated alongside a prohibition on the very activation operations it required. The executor subsequently cited that prohibition and a separate missing gate. The environment therefore contributes real instruction conflicts; these must not be attributed solely to Astra.

These cases support the user's complaint about operational reliability, while leaving the relative contributions of model, harness and prompt policy unresolved. The user's perceived decline since GPT-5.6 is a longitudinal experience report. No controlled historical comparison with earlier models was run for this issue.

What steps can reproduce the bug?

This is an observed incident with trace evidence, not a claimed deterministic minimal reproduction:

  1. Run a long Desktop task with Astra and allow it to approach automatic compaction.
  2. After a completed answer about task A, submit a clearly different task B.
  3. If a turn is interrupted during submission/compaction, compare the task B message displayed in the UI with persisted user-message events.
  4. Inspect compacted.payload.replacement_history and the active request selected after resume.
  5. In this incident, task B was missing and the agent resumed task A after "It froze".

Suggested regression test: submit A, complete A, enqueue B, interrupt before/during compaction, resume. Assert that B is either durably delivered exactly once or visibly remains unsent. Do not silently display B as submitted while the resumed model only receives A.

What is the expected behavior?

The newest submitted user request must survive interruptions and compaction, with a consistent UI and persisted transcript. Resume must follow the active request. If submission failed, surface that failure. Historical agent-written handoffs must not override original user instructions or be reported as verified facts without checking them. Corrections and status questions should steer authorized work without silently replacing it.

Additional information

Related: #27731 (stale prompt replay), #42782 (completed plan replay), #40646 (constraint drift). This report adds a specific UI-visible request absent from both the inspected user-message stream and replacement history, with interrupted-turn IDs and timings, rather than only a stale-prompt interpretation.

The user authorized filing this report. The public report contains reviewed excerpts and metadata. Full local rollouts and Telegram exports contain private operational information and have not been uploaded. Local evidence manifests retain source hashes and line references. The screenshot was examined locally and its relevant sequence transcribed above. No claim is made that private diagnostic logs were submitted through the in-app feedback channel.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the affected session's persisted user-message events and the compacted.payload.replacement_history, comparing them with the Desktop UI request around the two interrupted turns. Reproduce the A-then-B interruption and compaction sequence, then verify that the newest request is delivered exactly once or remains visibly unsent, and that resume follows it.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
desktop
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.