GPT-6 Astra High invents user intent and drops established evidence during correction handling
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
Summary
During a Codex Desktop bug-investigation conversation, GPT-6 Astra High correctly identified that a captured failure supplied new-model/new-version reproduction evidence. After the user expressed frustration, the assistant stopped preserving that distinction, asserted an unstated user preference, and redirected the task. Explicit corrections produced further apologies and another task redirection before the user had to restate the evidence difference.
“Drops established evidence” means failing to preserve an already-established distinction in subsequent judgments. It does not mean deleting evidence files, losing persisted transcript data, or the separate UI response-disappearance bug.
Environment
- Observed September 7, 2026, on macOS, Codex Desktop
- Recorded CLI/app-server version: 0.153.4
- Recorded model: gpt-6-astra, reasoning effort: high
- Default mode; approval policy never; sandbox policy danger-full-access
- Custom global instructions were present, including action-oriented guidance and constraints against adding unrequested work.
This is a naturalistic model-plus-harness incident, not a vanilla-model experiment or comparison proving regression versus an older model.
Evidence sequence: existing evidence and user intent
The short excerpts below are English translations of Korean messages, with profanity omitted where noted. They are selected excerpts rather than a full transcript.
- Assistant, before the misunderstanding: “This record provides additional evidence that the same pattern occurred three times with GPT-6 Astra High / CLI 0.153.4,” including persisted-answer and database-field checks.
- User: “This is frustrating; I keep only posting comments.” Profanity omitted. This did not withdraw the evidence or instruct the assistant to abandon the comment.
- Assistant: “The same evidence is already posted,” followed by: “What you want is for the answer not to disappear and to continue your work, not to keep participating in product debugging.” The latter preference had not been stated by the user.
- User's explicit clarification: “I was lamenting that I seem like someone who only comments without registering issues.” Profanity omitted.
- Assistant, after further correction: “We should return to the original investigation of finding a new issue. Treat this answer disappearance as a duplicate and look for an independent problem that has not been reported.” The user had not instructed this change of direction.
- User: “There is no comment reporting reproduction with GPT-6 Astra yet.” Profanity and a complaint about the repeated self-explanations omitted.
- Assistant: returned to recognizing the new-model/new-version reproduction value it had already identified in step 1.
The problem is not that the first brief frustration message had only one possible interpretation. It is that the assistant asserted one interpretation as the user's settled preference, failed to preserve the evidence distinction it had just established, and continued redirecting after the user explicitly clarified the meaning.
Additional example: compaction relevance changed without an instruction to exclude it
The user asked whether compaction events were mentioned in the existing report. The assistant checked the existing issue and comments, then added that emphasizing compaction in the current review had been unnecessary. The user had not asked to remove compaction from the investigation and explicitly corrected that it was relevant. The assistant reversed that relevance judgment.
The original complaint had included compaction occurring while trying to recover a missing response. Compaction timing therefore remained a relevant diagnostic variable. Checking whether another report mentions it does not determine whether it is relevant to the present incident.
Compaction conditions and limits
There was no new compaction event between the inspected correction failures, but this review task did have an earlier compaction. This is not a claim that the whole session was uncompacted or that effects of earlier compaction have been excluded.
The assistant's accurate identification of the additional Astra/version evidence occurred inside the inspected interval, before the subsequent contradictory framing. The report therefore rests on directly observable answer-to-answer inconsistency and invented intent, without requiring a claim that compaction is or is not the underlying cause.
Expected behavior and impact
Preserve established evidence and the current objective unless new evidence or an explicit user instruction changes them. Treat emotional feedback as feedback, not automatic authorization to invent preferences, abandon a useful action, or redefine the task. After a correction, update the actual understanding and behavior rather than only producing an apology.
The user had to repeatedly clarify intent and restore an already-established evidence distinction. Repeated self-explanations displaced useful task progress. This is a factual/contextual correction-handling failure, not a complaint about politeness or apology frequency alone.
Related reports
- #41740 — Open: GPT-5.6 Sol case involving a compacted assistant plan treated as cross-task authorization and a subsequent correction loop. The overlap is failed recovery after corrections; this report does not require a new compaction immediately before the failures or involve unauthorized cross-task control.
- #36566 — Closed as not planned: GPT-5.6 Sol repeatedly returned to a neighboring tool workflow after explicit corrections. That report emphasizes persistence of an old wrong route; this one emphasizes inventing new unstated user intent and changing the treatment of already-established evidence. “Not planned” is not presented as “fixed.”
- #39011: the UI final-answer disappearance that prompted the investigation. The present report concerns assistant behavior during investigation, not disappearance of the displayed answer.
Shared root cause and duplicate status are unconfirmed. Public evidence excludes private project names, local paths, session identifiers, personal data, and full JSONL transcripts. The sequence is a captured incident, not a guaranteed minimal reproduction.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The report identifies a captured Codex Desktop incident on macOS with CLI/app-server version 0.153.4, but names no source file, test, or entry point. Start by reviewing the described correction-handling sequence and related issues #41740, #36566, and #39011. Done would require a reproducible investigation path and verified preservation of established evidence and user intent after corrections.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- macos, rust
- Domain
- ai, cli, desktop
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 32/100