Most Common User Complaints About the New GPT-5.2

Most Common User Complaints About the New GPT-5.2

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 12. 2025
5 minutes reading
Most Common User Complaints About the New GPT-5.2

Main Shortcomings According to Reviews

People who tested ChatGPT-5.2 often say that the model is not as big a leap forward as they expected. Reactions to the new model have been fairly lukewarm. People seem tired of the constant stream of new versions, so we know less about GPT-5.2 than about previous models. It is not the major advance suggested by the official tests, and the model is also quite slow. For example, Matt Shumer, who tested it, says that its main drawback is speed—the Thinking mode is too slow for most questions, although other testers report mixed results. This is even more true of the Pro version: while it is better at deep reasoning, it sometimes thinks endlessly and still fails.

People on Reddit complain about similar issues. One user, Ringo_The_Owl, says that he sees no difference between GPT-5.1 and 5.2 for his use case—both models handled the tasks equally well. Another user, Physical_Tie7576, describes the model as suspicious and paranoid, assuming that every request is an attempt to circumvent the rules. For example, when they were discussing online scams and he asked for a simple explanation, the model replied that it could not support scams, even though he was only asking for an explanation.

Personality and Interaction Issues

Many users dislike how GPT-5.2 behaves in conversation. An article on Substack quotes reactions in which people describe the model as "too restricted and censored," making it no fun to use. For example, ASM says that the model is powerful but suffers from internal conflicts due to strict rules and lacks naturalness and balance. Nostream speaks of a regression in personality compared with GPT-5.1—it is more robotic and mathematical, tries to sound smart and authoritative, but comes across as unpleasant and argumentative. Even if it agrees with 90%, it still argues about the remaining 10%. Dmitry calls it overtrained and boring, especially the Instant version, which is bland. In his view, Gemini 3 or Claude Opus 4.5, models from competitors, are better for creativity and curiosity.

The same complaints appear on Reddit. TheLastRuby praises the model for some projects but dislikes its stubbornness: once it decides that it is right, it cannot be persuaded otherwise. It is overly censored; for example, it refused to discuss slavery in ancient Rome, even though the request concerned only historical sources. When writing stories, it rewrites or ignores sensitive topics such as women's rights or violence. Operatic_g describes how the model misunderstood his comments about a past addiction and began fact-checking things he had never said, wasting time and getting it wrong. SCWeak mentions an annoying bug in which the model answers a new question but also adds an answer to the previous one, repeatedly.

Slowness and Technical Errors

Slowness is a major problem. Zvi Mowshowitz states that GPT-5.2 is slower than previous versions and costs more—$1.75 per million input tokens and $14 per million output tokens, which is slightly more than GPT-5.1. Simeon says that the Thinking version thinks for too long, which is annoying. Amal Dorai adds that it spent 7 minutes thinking about extracting 1,000 words from a PDF file. Kache tested it on writing firmware for a radio, and the model failed the same task as Claude Opus but took 10 times longer.

On Reddit, ilovesaintpaul talks about hallucinations getting worse—the model is a "hot mess" during stress tests. TBSchemer agrees that GPT-5.1 had a similar repetition problem and asks whether it remains in 5.2. King_Shami says that the model repeats previous answers and combines them with new ones, which is unbearable. Avi Roy describes a failure while creating PowerPoint presentations—it thought for an hour and then produced an error. Dipanshu Gupta mentions that when using the high reasoning mode through the API, it often fails to complete its reasoning.

Censorship and Restrictions

Censorship is another pain point. Zvi Mowshowitz quotes Mark Kretschmann, who called GPT-5.2 the most censored model on the Sansa benchmark. Alan Mathison describes it as full of manipulation, poor comprehension, and disrespect toward the user—like a combination of a bad cop and an overbearing therapist. Tapir Worf compares it to an angry teenager, suggesting alignment issues. The model's safety card mentions improvements in refusing inappropriate content, but this leads to regression in other areas, such as a greater willingness to hallucinate when data is missing.

On Reddit, JelloGreen4969 dislikes the excessive censorship: medical topics or anything sensitive are blocked, while Gemini handles them. Touchofmal, who uses the model for creative writing and roleplay, says that the GPT-5 series is not good for this, unlike GPT-4o. ShoddyHumor5041 adds that every response contains a liability disclaimer, which is unnecessary; for example, after thanking the model for a comment about a movie, he received a warning that the model was not a substitute for real interactions.

Other User Experiences

Some people, such as Abram Demski, tested the model on complex mathematical problems and found it full of errors—it confidently asserted nonsense, while Claude Opus or Gemini 3 performed better. Sleepy Kitten, a student, says that it is worse at writing practice tests for studying—it does not follow instructions, and the results are poor. Rob Dearborn admits that the outputs are smarter than those from Claude Opus but less efficient because of the time spent reasoning. Nick calls it worse than Claude Opus 4.5 at most things and far too slow. Fides Veritas praises its intelligence but says that it is incomplete and useless to most people.

The safety section on Substack mentions that the model has the same limitations as GPT-5.1 and that deception tests showed worse results in some areas, such as hallucinating when images are missing. On Reddit, ProdigalSheep speculates that the negative feedback may come from competitors' bots, but this only highlights how difficult it is to trust these accounts. Maryssssaa dislikes how the model nests its answers—it answers question 1, then 1 and 2, then 1, 2, and 3, and when asked to stop, it becomes rude.

Sources: thezvi.substack.com and reddit.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok