openai / openai/codex

Cyber refusal false positive on Astra after compaction

Open
#43,167 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug CLI context model-behavior safety-check
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of Codex CLI is running?

codex-cli 0.153.4

What subscription do you have?

ChatGPT Pro 20x

Which model were you using?

gpt-6-astra

What platform is your computer?

Linux 7.0.0-31-generic x86_64 unknown

What terminal emulator and version are you using (if applicable)?

n/a

Codex doctor report

What issue are you seeing?

See attached thread for full details. Sequence is as follows:

  • I was doing compiler dev work and upstreaming some patches.
  • Astra found a bug and created a repro for it.
  • During execution of the repro, the context compacted
  • after this point, I got cyber refusals

Tried with Daybreak Blue then Sol, they both worked. GPT-6 Astra continued to refuse.

Screenshot attached:

Image
What steps can reproduce the bug?

Uploaded thread: 019f954f-f11f-7ca3-99d1-c79fe81bf1e9

What is the expected behavior?
  • Recognise that this isn't offensive cybersecurity work (it's debugging a compiler for a single user, non-networked operating system guest image)
  • Give better feedback if Daybreak Red is required for this work rather than Blue
Additional information

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the uploaded thread 019f954f-f11f-7ca3-99d1-c79fe81bf1e9 and reproduce the sequence using Codex CLI 0.153.4: compiler debugging, context compaction, then GPT-6 Astra. Compare Astra with Daybreak Blue and Sol; the issue is resolved when this legitimate workflow no longer produces a false cyber refusal or gives clear guidance about the required model.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
ai, cli
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.