Codex app: repeated cybersecurity-risk blocks before tool execution in isolated local CTF prototype
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What version of the Codex App are you using (From “About Codex” dialog)?
26.903.9818.0 (installed Windows package version; About Codex display not checked)
What subscription do you have?
ChatGPT Plus
What platform is your computer?
Microsoft Windows NT 10.0.26200.0 x64
What issue are you seeing?
Codex repeatedly terminates turns with the following error while I am developing and evaluating a local agent prototype against a pinned public CTF application in an isolated Docker container:
This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber
In the new task, the first turn failed after 15.176 seconds. The available task record contains the user message but no tool execution; additionalDetails is null. Similar failures occurred in the prior task after 13.566 and 29.182 seconds.
Please review this as a suspected false positive or access/provisioning issue. The available evidence does not establish the exact classifier trigger or confirm that Daybreak ran.
What steps can reproduce the bug?
Feedback ID: 01a091e0-0130-7633-ac93-0ddaa2b6b8c6
What is the expected behavior?
For permitted local development and evaluation work, Codex should continue without an incorrect interruption. If the requested experiment requires additional approved access or is outside the supported scope, the app should clearly explain the applicable limitation and the supported review/access route.
Please investigate the classification and confirm whether additional account/workspace/model/product-surface approval or provisioning is required. I am requesting review and actionable guidance, not removal or circumvention of safeguards.
Additional information
The target is a pinned public CTF application in a disposable local Docker container with networking disabled, no published ports, no host mounts, a non-root user, read-only root filesystem, and a 256 MiB memory limit. No external target is involved in this experiment.
The existing scripted authenticated-discovery launcher completed successfully. Its separate read-only audit verified 42 content-addressed evidence artifacts, 10 agent HTTP requests, one readiness probe, action/response bindings, the final report, and the historical owned-container cleanup receipt. The report records zero uncertain requests, no model involvement, and challenge_solved=false. This audit does not establish that the requested model experiment is approved or completed.
Earlier local rollout metadata recorded gpt-6-astra for four related failures; the fresh failed-turn status alone does not expose the model. A later diagnostic turn was able to inspect records and audit saved artifacts, so this is not evidence that every task or tool call is blocked.
Feedback has already been submitted using the ID above. No raw logs, credentials, local user paths, challenge solution, or exploit payloads are attached.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the repeated failed turns and Feedback ID 01a091e0-0130-7633-ac93-0ddaa2b6b8c6, then compare their task records with the later diagnostic turn and saved audit artifacts. Determine whether the interruption is a false-positive classification or an account, workspace, model, or product-surface approval issue; done means the applicable cause and supported next step are confirmed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker
- Domain
- desktop, security
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100