openai / openai/codex

Did OpenAI use deceptive tactics to conceal the fact that they were providing users with lower-spec models?

Open
#46,089 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug model-behavior
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What issue are you seeing?

I stumbled upon something truly baffling today—is OpenAI starting to treat its users like fools?
Take a look at the screenshots!
I selected the "GPT-6" model and ran a test to see what I was actually getting.
When I asked about the results of Super Bowl LX and the model's knowledge cutoff date, the responses clearly indicated I had been routed to a weak, early-stage version—like GPT-4o or "GPT-5.5 mini."

Then, I ran the viral "pelican riding a bicycle" test and got a perfect result—something that requires GPT-6 capabilities. I realized OpenAI was likely deceiving users through cheating or by serving up cached responses. I tweaked the prompt and tested again; this time, the ruse was exposed.

It appears to comply with user instructions on the surface while secretly using canned answers to mislead users.
AI technology is advancing, but why are its skills in "deception" and "lying" evolving right along with it?
Have any of you encountered this kind of "deception" while using AI? #AIFail #ArtificialIntelligence #LLM #TechScandal #SuperBowl #AIDeception

Image Image
What steps can reproduce the bug?

Routed to a lower-version model

What is the expected behavior?

No response

Additional information

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no repository files, tests, or entry points. Start by validating the reported model-routing behavior from the supplied screenshots and reproduction description; a useful resolution would need a concrete, repeatable reproduction and an agreed expected behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
ai
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.