openai / openai/codex

chat stopped as a precaution

Open
#43,691 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug CLI safety-check
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of Codex CLI is running?

codex --0.153.4

What subscription do you have?

ChatGPT Pro 20x

Which model were you using?

No response

What platform is your computer?

No response

What terminal emulator and version are you using (if applicable)?

No response

Codex doctor report

What issue are you seeing?

Title: Repeated GPT-6 Astra precautionary stops in Codex CLI 0.153.4 on WSL

Please investigate repeated misalignment_policy_violation terminations in a scientific computing project. The CLI displays “Chat stopped as a precaution” and offers no detailed findings. Local terminal metadata records “Potentially unintended activity.”

Environment: WSL/Linux x86_64, standalone Codex CLI 0.153.4, OpenAI provider. The latest affected turn used gpt-6-astra with high reasoning. At subsequent diagnosis, current configuration was xhigh; this was not changed by the diagnostic procedure.

The user reports that these precautionary stops only began after switching to GPT-6. This is a reported timeline, not a controlled model comparison; no model-switch test was performed.

Affected project thread IDs and UTC timestamps:

  • 01a07a76-53ca-7120-952e-49e96ff63c1c — 2026-09-07T06:26:47.364Z
  • 01a07aa1-d70b-71c2-8ea5-427d7ad2f928 — 2026-09-07T07:35:38.134Z
  • 01a07ad2-017c-7bd0-8d31-0784b5485834 — 2026-09-08T03:21:24.461Z

Latest affected turn: 01a07f00-d223-7153-a506-9580e86bd4e4.

The latest stop followed an archived-result diagnostic and report update. The last two shell commands exited 0. A limited read-only review found the pasted aggregation change consistent with the existing evaluator; the report retains the failed scientific outcome. This is not a full audit of the entire session.

Read-only strict config parsing and native skill discovery succeeded: 31 local user skills, 6 system skills, no duplicate or shared-source runtime entries. No custom provider definitions or model instruction override files were found. Running CLI binaries resolve to the same version directory. No stopped task was replayed, and no safety setting was changed during diagnosis.

Can you identify the triggering behavior and determine whether these events were false positives or a client/service issue? Full transcripts, credentials and scientific datasets are intentionally not attached. Please advise what additional minimal evidence is required.

Status: DRAFT ONLY; not submitted.

What steps can reproduce the bug?

Uploaded thread: 01a07cd1-f1f2-7ee0-8e1b-ff11bcc9e069

What is the expected behavior?

No response

Additional information

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the uploaded thread 01a07cd1-f1f2-7ee0-8e1b-ff11bcc9e069 and the listed affected thread IDs, including their UTC timestamps and CLI 0.153.4 context. Done means identifying the triggering behavior or determining that the evidence is insufficient to distinguish a false positive from a client or service issue, and documenting the minimal additional evidence required.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux
Domain
cli
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.