NVIDIA / NVIDIA/OpenShell

bug(gator): Codex auth refresh failures cause missing-sentinel retries

オープン
#2,467 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

state:stale
主要言語
Rust
スター
8.7k
フォーク
1.3k
平均マージ
2日 7時間
マージ済み PR(30日)
243

説明

Agent Diagnostic

  • Skills used: create-github-issue; gator debugging was performed from the local launcher/sandbox logs rather than GitHub checks.
  • OpenShell version tested: current local CLI reports openshell 0.0.88. The observed gator launch used the local openshell gateway, which earlier reported 0.0.75 during preflight before main was updated.
  • Current branch state checked: local main is clean and tracking origin/main.
  • Known fixes reviewed: current main includes the gator payload readability fix in the Docker build path, but the Codex harness still writes fallback/placeholder auth material when provider credentials do not include complete Codex auth fields.
  • Possible duplicate issues checked with gh api: no open issues matched gator Codex refresh, CODEX_AUTH_ACCESS_TOKEN refresh, or OPENSHELL_AGENT_RESULT token_expired.
  • Findings:
    • A supervised gator run against PR #2421 had early Codex cycles fail with websocket 401 Unauthorized followed by Failed to refresh token: 401 Unauthorized / token_expired.
    • The harness exited without emitting OPENSHELL_AGENT_RESULT, so the wrapper treated the cycle as a transient upstream transport failure and retried.
    • Later cycles recovered and completed successfully (merge_ready, then pr_merged), so the immediate PR watcher did not remain stuck.
    • scripts/agents/runtime/harnesses/codex/exec.sh creates /sandbox/home/.codex/auth.json, but when CODEX_AUTH_ID_TOKEN or CODEX_AUTH_REFRESH_TOKEN are absent it currently falls back to synthetic/placeholder values.
    • scripts/agents/gator/agent.yaml and scripts/agents/gator/providers/codex-gator.yaml provide access token/account ID fields but do not provide CODEX_AUTH_ID_TOKEN; refresh material is gateway-managed rather than placed in the sandbox.
  • Reason for filing: the run eventually succeeded, but the current behavior makes long-lived supervised gator watchers vulnerable to intermittent or expired Codex auth state. A few bad refresh attempts can consume retries, delay review cycles, or leave the failure looking like a generic missing-sentinel harness crash.

Description

Supervised gator runs that use the Codex harness can fail early cycles when the Codex access token expires or cannot be refreshed from the auth material written into the sandbox. In the observed PR #2421 run, Codex failed to connect to the websocket with 401 Unauthorized, then failed refresh with token_expired, and exited before writing the required OPENSHELL_AGENT_RESULT sentinel.

The supervisor wrapper did detect this as a transient upstream transport failure and retry. That was enough for this specific run to recover, post merge_ready, and later complete after the PR merged. However, the behavior is still fragile: the sandbox receives enough auth material to start Codex, but not necessarily enough valid material for Codex's own refresh path, while the gateway owns provider refresh outside the already-running process environment.

Expected behavior:

  • Gator should either receive complete, usable Codex auth material for the harness lifecycle, or refresh Codex auth through a gateway-mediated mechanism before each Codex cycle.
  • Missing unsupported auth fields should fail fast with an actionable preflight error instead of writing synthetic values into auth.json.
  • If a transient upstream auth/transport failure does happen, the wrapper should continue to emit a clear structured/sentinel result so the failure mode is not indistinguishable from a generic harness crash.

Reproduction Steps

  1. Configure a local openshell gateway with the codex-gator provider and run gator through the Codex harness.

  2. Launch a supervised PR watcher, for example:

    ./scripts/agents/run.sh --agent gator --gateway openshell --name gator-pr-2421-supervised --watch --background "Review and monitor PR #2421 through the gator-gate workflow. Scope this invocation only to PR #2421."
    
  3. Inspect the sandbox/launcher log at scripts/agents/gator/logs/<name>.log.

  4. When Codex auth is stale, the cycle fails with websocket 401 Unauthorized, then refresh token_expired, then exits without OPENSHELL_AGENT_RESULT.

  5. The wrapper retries the cycle as a transient watch failure. In the PR #2421 run, later cycles recovered and completed.

Environment

  • OS: Linux local development workstation.
  • OpenShell CLI: openshell 0.0.88 on current local main.
  • Gateway: local openshell gateway at https://127.0.0.1:17670 using mTLS.
  • Agent: scripts/agents/gator.
  • Harness: Codex, model gpt-5.5, reasoning high.
  • Latest release checked: not checked against public releases in this session; investigation was against updated origin/main.
  • Possible duplicate issues checked: yes, no open matching issues found.

Logs

Sanitized excerpts from scripts/agents/gator/logs/gator-pr-2421-supervised.log:

ERROR codex_api::endpoint::responses_websocket: failed to connect to websocket: HTTP error: 401 Unauthorized, url: wss://chatgpt.com/backend-api/codex/responses
ERROR codex_login::auth::manager: Failed to refresh token: 401 Unauthorized: {
  "detail": {
    "message": "Your authentication token has expired. Please sign in again to continue.",
    "code": "token_expired"
  }
}
ERROR: Your access token could not be refreshed. Please log out and sign in again.
openshell-agent: transient watch failure 1 (missing OPENSHELL_AGENT_RESULT after harness exit 1; upstream transport failure detected); retrying in 30s

The same supervised run later recovered:

OPENSHELL_AGENT_RESULT {"status":"waiting","next_poll_seconds":900,"reason":"merge_ready"}
OPENSHELL_AGENT_RESULT {"status":"complete","reason":"pr_merged"}

Suggested Direction

  • Extend the codex-gator provider schema and gator agent manifest to include the real CODEX_AUTH_ID_TOKEN if Codex requires it for stable operation.
  • Stop fabricating fallback ID tokens in scripts/agents/runtime/harnesses/codex/exec.sh; fail preflight with a clear message when required Codex auth fields are unavailable.
  • Keep the real refresh token gateway-owned. Do not copy it into the sandbox unless there is a deliberate security decision to do so.
  • Add a gateway-brokered refresh hook for supervised agents, such as a small openshell agent provider-env refresh helper, so long-lived watcher loops can request fresh provider env before each Codex invocation without exposing refresh tokens inside the sandbox.
  • Add a regression test or smoke path that verifies stale/expired Codex auth produces an actionable structured failure rather than a missing-sentinel harness exit.

Checklist

  • I have searched for existing related issues.
  • I included an agent diagnostic section.
  • I included the relevant sanitized logs.
  • I described the expected behavior and suggested implementation direction.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

scripts/agents/runtime/harnesses/codex/exec.sh から始め、次に scripts/agents/gator/agent.yaml と scripts/agents/gator/providers/codex-gator.yaml を調べて、sandbox に渡される認証フィールドを確認します。OPENSHELL_AGENT_RESULT がない場合の supervised watcher の動作と、サニタイズされたログを確認します。古い、または不完全な Codex 認証によって、対処可能な構造化エラーまたはサポートされている更新パスが生成され、回帰テストまたは smoke パスがある状態を完了とします。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
rust, shell, yaml
領域
authentication, devops, security
issue の種類
バグ
難易度
5/5
見積もり時間
1週間以上
活発さ
静か
明瞭さ
おおむね明確
初心者へのやさしさ
42/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。