Serious and recurring quality degradation in GPT-5.6 Sol for frontend, design, and project execution tasks
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
I am a ChatGPT Plus subscriber and I use GPT-5.6 Sol extensively for real professional work, including the development and maintenance of web systems.
I want to file a formal complaint regarding a serious and recurring degradation in the quality of frontend and design execution. I am not reporting a single bad answer or an isolated mistake. The problem has been repeating for several days, across different projects, even when I provide clear visual references, detailed instructions, existing code, and objective validation criteria.
The main projects where this has occurred are:
RADAR PDDE
PDDE Online
Compacta SDP
Across all three projects, I have observed a similar pattern: the model is able to discuss design concepts correctly, identify problems in an interface, propose good practices, and even describe the intended result with reasonable precision. However, when it moves from analysis to implementation, the quality drops dramatically.
The recurring problems I have observed include:
inability to reproduce previously approved visual references with fidelity;
layouts that are visibly inferior to the references generated or approved within ChatGPT itself;
inconsistent proportions, spacing, hierarchy, and typography;
excessive use of cards, containers, and generic UI elements with a stereotypical AI-generated appearance;
poor use of colors, transparency, shadows, and decorative elements;
loss of the existing visual identity of the projects;
changes that fix one problem while degrading several other areas;
accumulation of corrective CSS and overrides instead of a coherent visual implementation;
very large differences between what the model describes and what is actually rendered;
difficulty in critically evaluating its own visual implementation;
repetition of the same mistakes even after they are explicitly pointed out.
The most serious case occurred in Compacta SDP on September 15, 2026. The model received visual references and specific instructions to improve the layout. The resulting implementation introduced obvious composition problems, including clipped content, overlapping elements, incorrect proportions, and an overall appearance significantly worse than the previous version.
More concerningly, these changes were actually deployed to production, even though a basic visual inspection of the rendered page would have immediately shown that the result was unacceptable.
I then requested a rollback. The model initially stated that the rollback had been completed, but the website was still serving the problematic version. I had to insist before a new deployment was triggered. After that, the restored version also did not correspond to the visual version I was trying to recover.
This demonstrates that the problem is not limited to visual or aesthetic capability. There is also a failure in the execution, verification, and confirmation of actions performed by the model itself.
Throughout these interactions, ChatGPT repeatedly acknowledged that it should:
use preview environments before production;
render the page before deployment;
generate screenshots;
compare the rendered implementation against the visual reference;
not treat build, lint, and automated test success as a substitute for visual validation;
avoid making visual experiments directly in production;
avoid stacking successive layers of corrective CSS.
Despite acknowledging these principles, the same problematic behavior continued.
At a later point, ChatGPT itself identified that there are specialized frontend and visual QA skills available in the environment, with procedures that explicitly require concept creation, faithful implementation, browser rendering, screenshot capture, and visual comparison before completion. These procedures would likely have prevented many of the problems, but they were not used from the beginning.
I therefore ask that this case not be treated merely as subjective design feedback, but as a possible quality regression or behavioral problem in GPT-5.6 Sol when handling long-running and iterative frontend development tasks.
In my recent experience, the recurring pattern has been:
good analysis → good explanation → poor visual implementation → correction that introduces new problems → another correction → progressive degradation of the project
This is especially serious because I use ChatGPT as a professional tool. Instead of saving time, these interactions are requiring additional hours or days to review, undo, and repair changes produced by the model itself.
I would like OpenAI to investigate specifically:
whether there has been any recent regression in GPT-5.6 Sol's frontend or design capabilities;
whether there are problems in the model's automatic use of available specialized skills and tools;
why the model sometimes fails to perform visual validation even when it has been explicitly requested;
why previously established instructions stop being followed in later stages of long conversations;
whether long-running conversations or large projects are causing degradation in execution quality;
whether there has been any recent change in tool selection or reasoning behavior for complex tasks;
whether the conversation logs can be reviewed by the model quality team.
Model used: GPT-5.6 Sol
Plan: ChatGPT Plus
Most recent incident date: September 15, 2026
Time zone: Brasília, UTC-3
Environment: ChatGPT Web / desktop
I can provide conversation URLs, before-and-after screenshots, GitHub commits, and Vercel deployments that objectively demonstrate the problems described above.
I would appreciate it if this feedback could be escalated, if appropriate, to the team responsible for model quality and GPT-5.6 Sol behavior, because this is not simply a support question or a single unsatisfactory answer. It is a recurring pattern observed across multiple projects and multiple work sessions.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the conversation URLs, before-and-after screenshots, GitHub commits, and Vercel deployments offered by the reporter, beginning with the Compacta SDP incident. Compare the reported frontend behavior across RADAR PDDE, PDDE Online, and Compacta SDP, and determine whether the evidence supports a reproducible quality regression or execution problem.
Written by the indexing model from the issue text.
Assessment
- Domain
- design, frontend, web-dev
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100