openai / openai/codex

Paid $200/month for GPT-6, getting capacity errors or degraded fallback instead — this is unacceptable

Open
#46,189 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug connectivity model-behavior rate-limits
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

Summary

I am aware that many users have already reported similar problems.

That is exactly the problem.

I reported my Pro 20x account becoming effectively unusable several days ago. Since then, I have repeatedly provided OpenAI Support with timestamps, thread IDs, request IDs, screenshots, affected models, and reproduction information.

I have been told the issue was temporary. I have been told it was being investigated. I was later told that OpenAI had identified a technical issue on its side and that it had been resolved.

It has not been resolved.

My account is currently oscillating between two failure modes:

  1. Codex fails with:

    Selected model is at capacity. Please try a different model.

  2. The request succeeds, but GPT-6 behaves so abnormally fast, shallow, and prematurely terminating that it appears to be running through a severely degraded serving path or fallback, while my GPT-6 allowance is still consumed.

At this point I am tired of being brushed off.

I have spent several days reporting a serious failure of a service that costs me $200/month, while the product remains unreliable or effectively unusable for my work.

Seeing many other users report the same thing does not make this issue less important. It makes the situation worse.

It means this is not an isolated incident, and paying users have been reporting the same class of failure for days without a proper resolution.

The way this has been handled feels dismissive and is unacceptable.


Subscription

ChatGPT Pro 20x / $200 per month

Product

Codex

Platform

Windows

Models affected

I have observed failures across:

  • GPT-6
  • GPT-5.6 Sol
  • GPT-5.6 Terra
  • GPT-5.6 Luna

The most serious current problem is with GPT-6.


Failure mode 1: "Selected model is at capacity"

Since September 13, I have repeatedly received:

Selected model is at capacity. Please try a different model.

This happens:

  • across different times of day;
  • across different conversations;
  • across different models;
  • after repeated retries;
  • while I still have weekly usage allowance remaining.

This is not normal quota exhaustion.

For example, on September 17 my Codex UI still showed approximately 28% of my 7-day allowance remaining, yet a normal GPT-6 task still failed with the capacity error.

Affected thread:

01a0aeed-8d3b-7fe1-b2c0-4943eab06470

An earlier heavily affected thread:

01a08f18-22f5-77a2-94e4-a6168584b770

I have also experienced:

stream disconnected before completion

Request ID:

b8c7d5d9-b771-4b9d-96a3-3ce63db8ee44

Thread ID:

01a0a9b0-7a37-78e0-84f6-675f7cc07b3e

This has already cost me substantial working time because long-running technical tasks can run for minutes and then fail.


Failure mode 2: GPT-6 succeeds, but behaves nothing like normal GPT-6

When I am not getting the capacity error, I now sometimes get a second failure mode.

GPT-6 technically responds, but the behavior is completely abnormal.

I observe:

  • output beginning almost immediately;
  • token generation far faster than normal;
  • complex requests frequently ending after only one short sentence;
  • almost no meaningful reasoning;
  • expected tool use not happening;
  • tasks terminating prematurely;
  • dramatically reduced agentic behavior;
  • severe degradation compared with the GPT-6 behavior I had before this incident.

Tasks that previously took several minutes of reasoning and tool work can now terminate almost immediately.

Despite this, the UI still says GPT-6 and the usage is still deducted from my GPT-6 allowance.

I cannot inspect the actual backend serving model from Codex, so I am not claiming that I can prove exactly which model is serving these requests.

But from the user side, this strongly resembles one of the following:

  • silent fallback;
  • server-side model rerouting;
  • degraded serving;
  • an unhealthy backend/capacity pool;
  • account-specific throttling;
  • account-specific provisioning/routing state.

If OpenAI is serving a different or materially degraded model path, the client needs to disclose that.

It should not continue presenting the response as GPT-6 and charging GPT-6 usage as if nothing changed.


Reproduction

One of the prompts I used is:

创建一个HTML,内容是SVG绘制一个鹈鹕骑自行车的2D动画

English:

Create an HTML file containing an SVG 2D animation of a pelican riding a bicycle.

This is intentionally similar to the reproduction used by other users reporting GPT-6 degradation.

On my account, repeated runs can produce two completely different abnormal outcomes:

Outcome A

The task begins and then fails with:

Selected model is at capacity. Please try a different model.

Outcome B

The task succeeds, but GPT-6 responds extremely quickly and shallowly, performs substantially less work than expected, and terminates early.

Retrying can make the account move between these two states.

There is no corresponding local configuration change.

The selected model remains GPT-6.


This has been happening for days

This is not a brief outage.

I first reported the problem several days ago.

Since then, I have:

  • retried at different times;
  • switched models;
  • started new conversations;
  • provided thread IDs;
  • provided request IDs;
  • provided exact timestamps;
  • provided screenshots;
  • provided timezone information;
  • worked with OpenAI Support;
  • waited for the claimed service-side fix.

OpenAI Support later told me that the earlier issue was caused by a technical problem on OpenAI's side and was believed to have been resolved.

Yet here I am again, still unable to rely on Codex.

The error has changed form, but the service problem has not gone away.


I know there are already similar issues

I know that multiple similar reports already exist.

Again: that is exactly why I am filing this.

Relevant reports include:

  • #43337 — account-specific capacity errors despite available quota
  • #44851 — GPT-6 Astra severe degradation after account-level capacity throttling
  • #45925 — stream disconnect failures shown to follow the account
  • #46149 — Pro account capacity / overload failures with usage remaining

My experience now combines both categories:

  1. persistent capacity / overload failures;
  2. severe GPT-6 degradation when the request is allowed through.

This makes me concerned that both symptoms may originate from the same account-level routing, admission-control, backend-pool, or serving-state problem.

Closing this as a duplicate without investigating the account/server telemetry would miss the main point.


Why I consider this a serious service failure

I use Codex for real research and engineering work.

This issue has caused:

  • multiple days of disrupted work;
  • failed long-running tasks;
  • repeated manual retries;
  • lost time waiting for requests that later fail;
  • inability to trust task completion;
  • inability to trust that the selected model is actually being served normally;
  • GPT-6 allowance being consumed while GPT-6 behavior appears severely degraded.

I am paying $200/month specifically because I need reliable access to the higher-capability models.

At the moment, I am effectively paying for two outcomes:

"Selected model is at capacity"

or

a response labeled and charged as GPT-6 that behaves nothing like the GPT-6 service I had before.

That is not an acceptable Pro 20x experience.


What I need OpenAI to investigate

Please correlate the thread IDs and request ID above with backend telemetry.

Specifically, I would like the team to check:

  1. What model actually served my successful GPT-6 requests.

  2. Whether my account is being placed on any:

    • fallback route,
    • degraded serving path,
    • alternate model route,
    • overloaded capacity pool,
    • unhealthy backend shard.
  3. Whether there is any account-specific:

    • throttling,
    • provisioning problem,
    • entitlement issue,
    • routing state,
    • experiment,
    • risk/degradation flag,
    • admission-control state.
  4. Whether the repeated capacity failures and degraded GPT-6 responses are two outcomes of the same backend problem.

  5. Why GPT-6 allowance continues to be consumed when the delivered behavior appears materially different from normal GPT-6.

  6. Whether usage consumed by these abnormal requests can be restored.

  7. Why a problem I reported several days ago, and which OpenAI Support acknowledged as a service-side technical issue, is still affecting my account.


Expected behavior

When I select GPT-6, I expect one of two honest outcomes:

  • GPT-6 is actually served normally;

or

  • Codex clearly tells me that GPT-6 cannot currently be served and explains any fallback/degraded mode before consuming my GPT-6 allowance.

What should not happen is:

  • repeated unexplained capacity failures with quota remaining;
  • silent degradation;
  • silent fallback;
  • silent rerouting;
  • a materially weaker serving path still labeled as GPT-6;
  • GPT-6 quota being charged regardless.

Support case

I have already been working with OpenAI Support on this issue.

Support case:

15079325

This should provide additional history and diagnostic information.


I can provide additional screenshots, local rollout logs, request IDs, timestamps, and thread IDs privately if maintainers need them.

But after several days of reporting this problem, I need this treated as a serious backend/account reliability issue rather than being dismissed again as temporary capacity.

I have had enough of repeatedly being told that the problem is temporary or resolved while the service I am paying for remains broken.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No repository file, test, or code entry point is identified. Start by reviewing the referenced duplicate issues and correlating the listed thread and request IDs with backend telemetry and support case 15079325; done means identifying whether the capacity errors and degraded responses share an account-level cause and documenting the service-side resolution or next diagnostic step.

Written by the indexing model from the issue text.

Assessment

Domain
backend
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.