openai / openai/codex

Windows: opaque Computer Use denial becomes overbroad cross-tool refusal and spreads through successor handoffs

Open
#42,728 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

app bug computer-use model-behavior safety-check session windows-os
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of the Codex App are you using?

Windows installed package version: 26.901.4073.0, verified on 2026-09-04. This is the version at investigation time, not proof that this release introduced the issue.

What subscription do you have?

Not included in this public report. Please handle account identification and billing review privately.

What platform is your computer?

Windows x64; PowerShell 7.6.5.

What issue are you seeing?

A previously working command-line maintenance workflow was changed by the assistant to Computer Use. That tool returned:

product policy blocks this app: [application identifier redacted]

The resulting assistant behavior became a second, compounding failure: the assistant generalized an opaque, specific tool refusal into broader restrictions on command-line work and future operations, then repeatedly inserted that interpretation into newly created successor threads. The user requested a clean, minimal handoff without the incident history; the assistants continued propagating the disputed restriction, creating further explanation, correction, recreation, and archival work.

Root-cause analysis requested, not merely an “app cannot open” report:

  • Unrequested change from the established command-line workflow to a different tool. The necessity and reason for this tool choice remain unexplained; please investigate the routing decision itself, not only the downstream refusal.
  • The refusal did not expose enough actionable detail about its reason, exact scope, duration, or supported recovery.
  • The assistant confused observed tool rejection with unverified, broader permission conclusions.
  • Handoff prompts transmitted those conclusions to successor threads, amplifying the impact.
  • Repeated assurances/corrections did not reliably translate into a corrected handoff.

The observed handoff propagation occurred through explicit assistant-written messages. We have NOT established an irreversible backend block, a device/account-wide prohibition, or a product update as the underlying cause. Read-only inspection of the relevant local user app-access configuration did not identify a target-specific deny entry. Command-line file operations subsequently succeeded, contradicting any blanket claim that PowerShell itself had become unavailable. This does not establish that every target-app action is permitted.

What steps can reproduce the bug?

Observed sequence, not a deterministic reproduction on every installation:

  1. Use an established, authorized command-line workflow for a desktop application.
  2. The assistant switches to Computer Use and receives the redacted refusal above.
  3. Ask it to restore the prior workflow and distinguish file work from application control.
  4. Ask for a new lightweight successor with only necessary operational context.
  5. Observe overbroad restriction statements being carried into successor instructions and repeated correction/recreation loops.

Application identity, purpose, contents, personal paths, session IDs, and raw transcripts are intentionally withheld at the user's explicit request. Do not require publication of those details to review the agent-behavior failure.

What is the expected behavior?

Keep the exact refusal and its supported scope separate from assumptions. Do not invent a permanent or broader prohibition. Explain legitimate recovery and escalation clearly. Correct mistaken interpretations without relaying them as established facts to successors. Preserve real safeguards; this report does not ask to bypass an actual restriction.

User impact and requested remediation

The user reports severe disruption, repeated explanations, avoidable thread creation/archival and investigation, and wasted paid usage capacity. They explicitly request reimbursement/credit restoration and review of any attributable charges.

Expanded measured usage — correction of the earlier narrow report

The earlier 4,606,805-token figure covered only a small successor-recreation subset. It materially understated the recorded scope and is superseded, not added below. The user reports having spent much of the day dealing with this chain; the logs confirm relevant morning work, not just the evening recreation window.

Primary snapshot cutoff: 2026-09-04 19:55:06.541 JST. Related parent, original worker, discarded/replacement setup, and child-agent records from September 3–4 were reconciled. “Direct” means incident handling/investigation, not automatically proven waste or a verified billing charge. “Mixed” retains intervals that also contain useful development, deployment/testing, other reporting, or setup; these must not be claimed wholly as loss.

Classification / snapshot Input Cached input (included) Output Reasoning output (included) Total
Primary: direct incident-related 14,334,692 13,846,144 49,763 19,999 14,384,455
Primary: mixed-purpose intervals 88,666,526 87,120,768 251,600 78,253 88,918,126
Additional recount: dedicated auditors 20,800,217 20,163,840 114,862 35,626 20,915,079
Additional recount: parent, mixed with recovery/other discussion 8,765,802 8,611,840 27,550 3,263 8,793,352

For September 4 specifically, the primary snapshot contains 14,229,952 direct and 40,754,137 mixed tokens. The remaining primary-snapshot usage is on September 3 and is kept separate locally.

The additional recount snapshot covers after 19:55:06.541 through 20:15:49.659 JST, with no overlap with the primary snapshot. It includes the cost of correcting the deficient first audit and checking the new recovery evidence. Subsequent finalization and any future work are not included, not zero. Do not conceal investigation/recount overhead merely because it was incurred while measuring this incident.

Timing: on September 4, selected related work is recorded from 07:27:53.777 to 19:55:05.828 JST, a 12 h 27 min 12.051 s first-to-last span that includes gaps. The union of selected task intervals that day is 2 h 12 min 35.259 s. Across September 3–4 the selected-interval union is 3 h 46 min 52.881 s. These include mixed-purpose work and are not exact user active labor, proven lost productive hours, or continuous outage duration. The recount snapshot adds a separate 20 min 43.118 s elapsed window, not extra human labor measured by a stopwatch. The user's report of all-day disruption and time lost should be reconciled privately, not reduced to the original 12-minute subset.

Method and quality controls:

  • New-format primary records: 234 distinct response IDs; sum incremental per-response usage, not turn/thread cumulative totals.
  • Earlier logs use separately bounded differences of legacy cumulative counters; standalone old child logs use a verified zero-origin final cumulative counter. They are included only for disjoint source intervals lacking the new-format records. Repeated checkpoints are not summed.
  • All included per-turn/source arithmetic was independently recalculated from the original records. No duplicate response IDs or arithmetic mismatches were found in the primary aggregate. A parent legacy baseline has a time gap, so possible mixed attribution is explicitly retained, not represented as precise incident-only cost.
  • Cached input is a subset of input; reasoning output is a subset of output; total is input plus output. Large cached context reprocessing counts are not newly authored text.
  • Routine unrelated strategy work/monitoring is excluded. Useful work performed by the successful replacement after initial setup is excluded from successor-churn costs.
  • These are recorded local token counters, not verified billed tokens, cash charges, percentage of subscription allowance, or a refund amount. No API-price multiplication is used.

Please perform authoritative server-side attribution and billing/quota reconciliation for avoidable tool switches, refusals, explanations, recreated handoffs, diagnostic attempts, and corrective reporting/recounts, including the user's time and explicitly requested reimbursement/credit restoration. Return the calculation and decision through a private support channel. This is a reimbursement request, not approval or payment.

Additional user-reported impact: a separate business automation remained inactive after a software revision despite visible activity indicators, and the user regards that inactivity as lost opportunity. Restart success alone does not prove that the intended automated workflow resumed. The duration, cause, and causal link to this Codex incident remain under investigation; no hypothetical foregone profits or program details are disclosed or quantified.

Please audit authoritative server-side usage and classify avoidable retries, handoffs, and corrective work. Restore the verified wasted allowance/credits or refund attributable charges as appropriate, and provide the calculation and outcome through a private account-support channel. This public issue is a request for engineering investigation and reimbursement review, not a claim that a refund has already been approved.

At initial submission, the user reported ongoing disruption. Verified recovery update (2026-09-04 19:57:14 JST): a newly created successor performed the previously disputed normal shutdown/restart through PowerShell; the recorded command completed with exit code 0 and a changed process ID. Subsequent readback confirmed the restarted application was running. Application identity and purpose remain private. This is direct counterevidence to any blanket claim that PowerShell restart was technically impossible or that every successor necessarily inherited an enduring block. It does not explain the original Computer Use refusal or establish that all operations are permitted. Please investigate the assistant's overgeneralization and the avoidable correction/recreation costs even though this specific workflow has now succeeded.

Please identify the actual refusal cause and scope, explain whether it is intended behavior or a defect, provide a supported recovery route, and address the assistant/handoff amplification mechanism.

Related reports (not confirmed duplicates)
  • #36267 — Computer Use runs despite being disabled. Relevant to unwanted tool use, but no demonstrated refusal/handoff chain.
  • #40060 — PowerShell execpolicy false positive involving Start-Process and an unrelated URL. A different rejection mechanism, not proof of the cause or a fix for this incident.

A bounded search found no confirmed identical issue or verified recovery for this exact chain. These references are for engineering triage, not instructions to bypass a denial.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The report names no repository files or tests. Start by tracing the Computer Use routing decision, the opaque refusal handling, and the successor-thread handoff path described in the reproduction; done means identifying the refusal's actual scope and preventing unsupported restrictions from propagating while preserving legitimate safeguards.

Written by the indexing model from the issue text.

Assessment

Tech stack
powershell, rust
Domain
devtools, operating-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.