openai / openai/codex

Excessive usage consumption, poor task efficiency, and ineffective previous feedback

Open
#44,455 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug CLI model-behavior rate-limits
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of Codex CLI is running?

codex-cli 0.153.4

What subscription do you have?

ChatGPT Plus

Which model were you using?

gpt-5.6-sol medium

What platform is your computer?

Windows x64

What terminal emulator and version are you using (if applicable)?

No response

Codex doctor report

What issue are you seeing?

Codex is consuming a disproportionate amount of my 5-hour usage limit, sometimes even for very small tasks.

For example, my remaining 5-hour usage dropped from 91% to 89% after a very small task that only modified AGENTS.md with a few simple operational rules. No build, compilation, tests, or other files were executed or modified.

More importantly, this is not only a usage problem. In several previous development sessions, Codex consumed a very large percentage of my usage limit while producing incorrect or incomplete results. Because of those errors, I then had to repeat substantial parts of the work from scratch, consuming even more usage for work that had effectively already been paid for once.

This has happened repeatedly: Codex performs a long task, consumes a significant percentage of the available usage, but the result contains errors, does not follow the requested instructions, or cannot be used. The task then has to be corrected or restarted, sometimes almost entirely.

This makes the effective cost much higher than the percentage consumed by the initial task, because I am also paying in usage for correcting Codex's mistakes and repeating work.

I have submitted feedback through Codex several times in the past about poor results and problems during these sessions, but I did not realize that "Feedback uploaded" still required me to manually create a GitHub issue. Therefore, those previous reports were not accompanied by GitHub tickets.

My main concerns are:

  • usage consumption appears disproportionate to the actual work performed;
  • very high usage can be consumed even when the resulting work is incorrect or unusable;
  • Codex errors have repeatedly forced me to redo work from scratch;
  • correcting Codex's own mistakes consumes additional usage;
  • Codex has recently become significantly less efficient and reliable for my development workflow;
  • the feedback workflow is unclear because "Feedback uploaded" can give the impression that the report has already been submitted for investigation.

Uploaded thread ID for this report:
01a085dc-6810-7cb2-b509-61d96daa832c

What steps can reproduce the bug?

Uploaded thread: 01a085dc-6810-7cb2-b509-61d96daa832c

What is the expected behavior?

Usage should be reasonably proportional to the complexity and amount of useful work performed.

A very small task, such as modifying a few operational rules in a single file without running builds or tests, should not consume a significant portion of the 5-hour usage limit.

More importantly, Codex should reliably follow the provided instructions and project rules. When substantial usage is consumed, I expect the resulting work to be correct and usable.

If Codex produces incorrect or incomplete work, users should not repeatedly lose significant portions of their usage allowance simply to have Codex correct its own mistakes or redo the same work from scratch.

There should also be a way to restore or reset usage that was consumed because Codex produced unusable results or repeatedly failed to complete the requested task correctly.

I have also purchased €20 of additional Codex credits. Despite performing only a limited amount of useful work, 93 of those purchased credits have already been consumed. At this rate, using Codex for serious development work is not economically sustainable for me.

I would therefore expect OpenAI to investigate both the abnormal usage and, where excessive usage resulted from Codex failures or repeated rework, restore the affected usage/credits.

Usage should also be transparent enough for users to understand exactly why a task consumed a particular amount of their available limit.

Additional information

I am developing SARDU Edu as a free and open-source educational project based on open-source Scratch components.

The project is being developed with the intention of remaining freely available and respecting the licensing requirements of the upstream Scratch components. It is not being developed as a revenue-generating product.

For this reason, I have no business revenue from SARDU Edu that can absorb unpredictable or disproportionately high AI development costs. I am personally funding its development, including my ChatGPT Plus subscription and the additional Codex credits I have purchased.

I have already purchased €20 of additional credits, and 93 credits have been consumed after a relatively small amount of useful development work. When Codex produces incorrect results and the same work has to be repeated, these costs become particularly difficult to justify and unsustainable for a free educational project.

I am not asking for special treatment because the project is free. I am asking OpenAI to investigate the unusually high usage, including previous sessions where substantial usage was consumed for incorrect or unusable results, and to consider restoring usage or credits where excessive consumption or repeated rework can be confirmed.

I would like to continue developing SARDU Edu with Codex, but the development cost needs to remain reasonably proportional to the useful work performed.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the uploaded thread 01a085dc-6810-7cb2-b509-61d96daa832c and the reported usage changes for the small AGENTS.md task. Done means establishing whether the usage was abnormal, whether the reported failures are reproducible, and whether the “Feedback uploaded” workflow needs clarification.

Written by the indexing model from the issue text.

Assessment

Domain
cli
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.