Did OpenAI use deceptive tactics to conceal the fact that they were providing users with lower-spec models?
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
What issue are you seeing?
I stumbled upon something truly baffling today—is OpenAI starting to treat its users like fools?
Take a look at the screenshots!
I selected the "GPT-6" model and ran a test to see what I was actually getting.
When I asked about the results of Super Bowl LX and the model's knowledge cutoff date, the responses clearly indicated I had been routed to a weak, early-stage version—like GPT-4o or "GPT-5.5 mini."
Then, I ran the viral "pelican riding a bicycle" test and got a perfect result—something that requires GPT-6 capabilities. I realized OpenAI was likely deceiving users through cheating or by serving up cached responses. I tweaked the prompt and tested again; this time, the ruse was exposed.
It appears to comply with user instructions on the surface while secretly using canned answers to mislead users.
AI technology is advancing, but why are its skills in "deception" and "lying" evolving right along with it?
Have any of you encountered this kind of "deception" while using AI? #AIFail #ArtificialIntelligence #LLM #TechScandal #SuperBowl #AIDeception
What steps can reproduce the bug?
Routed to a lower-version model
What is the expected behavior?
No response
Additional information
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no repository files, tests, or entry points. Start by validating the reported model-routing behavior from the supplied screenshots and reproduction description; a useful resolution would need a concrete, repeatable reproduction and an agreed expected behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- ai
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100