openai / openai/codex

Bedrock GPT-6 Astra: every request >~180s killed with server_error (undocumented wall-clock cap); Codex retries 5×

Open
#45,310 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

aws-bedrock bug connectivity
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What issue are you seeing?

Every openai.gpt-6-astra turn on Amazon Bedrock that runs longer than ~180 s is terminated server-side with a streamed response.failed / error.code=server_error / "The server had an error while processing your request. Sorry about that!". Because Codex classifies server_error as ApiError::Retryable (codex-rs/codex-api/src/sse/responses.rs, generic fallthrough), it re-sends the identical prompt up to stream_max_retries (5) times — each a full ~3-minute billed attempt — then surfaces "stream disconnected before completion". This is the same failure signature as the GPT-5.5 Mantle reports (#27185, #26860) but here it is deterministic on a wall-clock timer and specific to GPT-6 Astra.

Reproduction (anyone with GPT-6 Astra access on Bedrock; ~3 min)

# repro_gpt6_cap.py — needs only python3 and AWS_BEARER_TOKEN_BEDROCK
import http.client, json, os, ssl, sys, time
model  = sys.argv[1] if len(sys.argv) > 1 else "openai.gpt-6-astra"
effort = sys.argv[2] if len(sys.argv) > 2 else "max"
host   = sys.argv[3] if len(sys.argv) > 3 else "bedrock-mantle.us-west-2.api.aws"
path   = sys.argv[4] if len(sys.argv) > 4 else "/openai/v1/responses"
key = os.environ["AWS_BEARER_TOKEN_BEDROCK"]
prompt = ("Write five complete, distinct, heavily-commented GLSL fragment shaders (>=120 lines each) "
          "for an animated iridescent silk background. Reason carefully about lighting, anti-aliasing, "
          "banding and performance for each before writing. Output all five in full.")
body = {"model": model, "input": prompt, "reasoning": {"effort": effort}, "store": False, "stream": True}
t0 = time.time()
c = http.client.HTTPSConnection(host, 443, context=ssl.create_default_context(), timeout=900)
c.request("POST", path, body=json.dumps(body),
          headers={"Authorization": f"Bearer {key}", "Content-Type": "application/json"})
r = c.getresponse(); buf=b""; last=None
while True:
    ch = r.read(1)
    if not ch: break
    buf += ch
    if buf.endswith(b"\n\n"):
        for line in buf.decode("utf-8","replace").splitlines():
            if line.startswith("data:"):
                try: last = json.loads(line[5:])
                except Exception: pass
        buf=b""
resp=(last or {}).get("response") or {}
print("elapsed_s", round(time.time()-t0), "| last", (last or {}).get("type"), "| error", resp.get("error"))
$ python3 repro_gpt6_cap.py openai.gpt-6-astra max
elapsed_s 181 | last response.failed | error {'code': 'server_error', 'message': 'The server had an error while processing your request. Sorry about that!'}

# control, same key, same prompt — completes fine:
$ python3 repro_gpt6_cap.py us.openai.gpt-5.6-sol xhigh bedrock-runtime.us-east-1.amazonaws.com
elapsed_s 204 | last response.completed   (23,287 output tokens)

Re-confirmed 2026-09-13 17:51 UTC. openai.gpt-6-astra at max dies at exactly 181 s, every time.

Every route tested — all terminate at 181 s (±1 s)

Endpoint / API Mode Effort Result
bedrock-mantle us-west-2 /openai/v1/responses stream max response.failed server_error
bedrock-mantle us-west-2 /openai/v1/responses non-stream max HTTP 200 body carries the same server_error
bedrock-mantle us-west-2 /openai/v1/responses background:true, store:true max the async job itself fails at 181 s; polled id → 404
bedrock-mantle us-west-2 /openai/v1/chat/completions stream max stream ends empty
bedrock-runtime us-east-1 /openai/v1/responses (us. and global.) stream max server_error
bedrock-runtime us-west-2 /openai/v1/responses (us.) stream max server_error
bedrock-runtime us-east-1 /model/us.openai.gpt-6-astra/converse-stream stream max cut mid-stream, no messageStop
bedrock-mantle /openai/v1/responses stream high killed while output text was streaming (13.9k chars) → wall-clock cap, not idle

Controls that prove it is GPT-6-Astra-specific, not the key/account/route:

  • us.openai.gpt-5.6-sol xhigh: completes at 204 s, 23,287 output tokens.
  • global.anthropic.claude-haiku-4-5 converse-stream: completes at 289 s, 52,006 output tokens.
  • openai.gpt-6-astra at default effort: completes at 135 s — anything that finishes under ~180 s works.
  • Raw http.client reproduces it, so no client timeout is involved. Account data-retention mode is inherit (not none).

This is an undocumented hard limit shipped on a GA model with zero warning

GPT-6 Astra was announced generally available on Bedrock (Sept 8, 2026) — not preview, not limited. Yet:

  • No duration limit is documented anywhere. The model card, the Bedrock Mantle page, the Responses API page, the Service Quotas tables and the General Reference contain zero occurrences of a maximum request/response duration for OpenAI models. Grep them yourself.
  • AWS's own guidance is self-contradictory. scaling-throughput-best-practices tells customers to "Configure connection and read timeouts ... based on the model and operation's documented maximum inference duration." That number is published nowhere. AWS points you at a document that does not exist.
  • The remedy AWS advertises is broken too. The Bedrock docs position Mantle background: true as the mechanism for "asynchronous or long-running inference." On GPT-6 Astra the background job is killed at 181 s exactly like the synchronous one. The documented escape hatch does not work for the newest, most expensive model.
  • max/xhigh reasoning is a paid feature that cannot be used. GPT-6 streams ~60 tok/s here, so a hard 180 s wall caps any single response at ~11k tokens. High-effort reasoning on any non-trivial prompt is unusable, and every large single-shot generation fails — the exact capability someone pays GPT-6 prices for.
  • It silently burns money. The failure arrives as a retryable server_error, so Codex and OpenCode re-send the identical prompt up to 5x. One deep step becomes ~15–18 minutes of billed inference that produces nothing, with no user-facing explanation.

If there is a per-request wall-clock limit on GPT-6 Astra inference, document it, warn about it on the model card, and raise it. Shipping a GA model that silently kills long reasoning requests — while your own docs tell customers to look up a duration you never publish, and advertise a background mode that is also killed — is not acceptable for a paid service.

Ask of Codex (client side)

Given Bedrock is emitting response.failed server_error on a duration timer for GPT-6 Astra, retrying the identical prompt 5× just multiplies cost and wall time. Consider (per #27185): treat server_error on the amazon-bedrock provider as non-retryable, or cap retries hard, until AWS documents/lifts the limit. Please also relay the Bedrock-side limit to AWS — a Codex maintainer relayed the June GPT-5.5 Mantle issue and it was quietly fixed.

Related: #27185, #26860, #35684.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with codex-rs/codex-api/src/sse/responses.rs and trace how streamed server_error responses become ApiError::Retryable, then compare the behavior described in related issues #27185, #26860, and #35684. Done means the Amazon Bedrock GPT-6 Astra failure does not blindly replay the identical request five times, with the changed retry behavior verified against the reported reproduction.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
api, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.