Don't stop here
Hand-picked guides our readers explore right after this one.
Master ChatGPT with advanced prompting techniques, mega-prompts, and proven frameworks
Read the guideResearch-grade prompts for Perplexity AI's search-powered responses
Read the guideAI prompts for idea generation, creative thinking, problem solving, and innovation
Read the guide'ChatGPT got worse' is one of the most searched complaints about the product, and it is a mix of something real and something perceptual. The real part: the GPT-5 line was tuned for different goals than GPT-4 era models, trading verbose, agreeable helpfulness for faster inference, tighter safety behaviour, stronger benchmark performance on reasoning and code, and shorter answers by default. Cost-driven routing compounds it, because the router sends most turns to a fast model and only escalates to reasoning when it judges the question hard, so identical prompts can get very different depth. Model retirements make it abrupt: GPT-4o was removed in early 2026 and GPT-5.2 followed on 12 June 2026, with conversations migrated to the matching GPT-5.5 model, so carefully tuned prompts changed behaviour overnight through no action of yours. The perceptual part: your expectations rose, your tasks got harder, and the novelty wore off. The good news is that most of the lost depth is recoverable through model selection and prompt structure rather than by cancelling.
Answers are noticeably shorter and more clipped than they used to be
It ignores parts of a multi-part instruction or drops constraints
More hedging, more refusals, and more 'consult a professional' framing
Output quality drops partway through a long working session
Prompts that reliably worked for months suddenly produce worse results
It summarises a task instead of actually doing it
The single biggest driver of inconsistency. A router picks between a fast instant model and a deeper reasoning model per turn, and it favours the cheap path. The model is not degraded so much as you are frequently not talking to the strong one.
GPT-4o was retired in early 2026 and GPT-5.2 was removed on 12 June 2026, with existing chats moved to GPT-5.5 equivalents. A prompt tuned against a retired model can behave noticeably differently on its successor with no change on your side.
Later releases were tuned toward tighter guardrails and less flattery, which shows up as more refusals, more caveats, and less personality. Users experience that as reduced helpfulness even where factual accuracy is unchanged.
The GPT-5 line answers more concisely unless you ask otherwise. If your old workflow relied on the model volunteering detail you did not explicitly request, it now looks lazy.
Deep in a long conversation, earlier context gets truncated and the model loses decisions made an hour ago. That reads as the model getting dumber over a session when it is really context loss.
When to try: First, on any task where quality matters
Select the Thinking option in the model picker rather than leaving routing to the system. Most complaints about shallow, careless answers come from being served the fast model on a task that needed the deep one. On Free and Go, Thinking sits behind the '+' menu in the composer and is capped.
When to try: Whenever answers feel thin
State the length, structure, and rigour: 'give me roughly 800 words, in five sections, with your assumptions stated and any uncertainty flagged'. Brevity is now the default, so depth has to be requested rather than assumed.
When to try: After any announced model retirement
Re-test your saved prompts, custom instructions, Projects instructions, and Custom GPT instructions against the current model. Restate output format, tone, and what 'done' looks like. Prompts tuned for a retired model are the most common cause of a sudden personal quality drop.
When to try: When quality drops partway through a long session
When a session degrades, open a new chat and paste in a short brief: the goal, the decisions already made, the constraints, and the current state. This fixes the truncation-driven 'it forgot everything' version of the problem, which no model choice can fix.
When to try: To stop re-specifying preferences every time
Use Settings, then Personalization, then Custom Instructions to state your defaults once: preferred depth, whether you want caveats trimmed, formatting, and that you want assumptions surfaced. This restores much of the lost 'personality' without repeating yourself each chat.
When to try: When answers are all caveat and no content
If you get a hedge or refusal on a legitimate task, restate the context and purpose plainly and ask for the specific deliverable rather than an open-ended request. Guardrail tuning fires on phrasing as much as on substance. Note that for medical, legal, and financial questions, caution is appropriate: treat any answer as general information, not professional advice.
When to try: On any factual or high-stakes output
Ask for sources, then check them. Ask the model to critique its own answer and list what could be wrong. This catches the confidently-wrong failure mode that makes a fast-routed answer feel worse than an old thorough one.
When to try: Before changing your subscription
Run the same prompt through Claude, Gemini, or Perplexity. If all three underperform, the prompt is the problem. If only ChatGPT does, you have a real basis for switching or for choosing the right tool per task rather than acting on a vague sense of decline.
Keep a small library of tested prompts and re-validate them after model retirements
Set depth, format, and tone once in custom instructions instead of per chat
Pick the reasoning model deliberately rather than trusting automatic routing
Start new chats per task so context truncation never masquerades as degradation
This is product feedback rather than a support case. Use the thumbs-down control on a specific bad response so it reaches OpenAI with the conversation attached, and include what you expected. Contact help.openai.com only if a paid feature is missing, the model picker fails to load, or your account shows limits below your plan. General 'quality dropped' tickets rarely resolve without concrete examples.
Both. Measurably, the GPT-5 line was tuned toward faster inference, tighter safety behaviour, and shorter answers, and cost-driven routing sends most turns to a lighter model. Perceptually, expectations rose and tasks got harder. It is better described as retuned for different goals than as dumber, and much of the lost depth is recoverable through model choice and prompting.
Model retirements. GPT-4o was removed in early 2026 and GPT-5.2 Instant, Thinking, and Pro were retired on 12 June 2026, with existing conversations migrated to the matching GPT-5.5 model. Prompts tuned against a retired model often need their output format, depth, and tone restated for the successor.
Select the Thinking model rather than leaving it on Auto, ask explicitly for the length and structure you want, put your standing preferences in custom instructions, and start a fresh chat with a short summary when a long session degrades. Those four changes recover most of the perceived loss.
Long threads exceed the usable context and earlier turns get truncated, so decisions made earlier vanish. It looks like the model deteriorating but it is context loss. Carry a written summary into a new chat instead of extending the same thread indefinitely.
Test before you decide. Run the same prompts through Claude, Gemini, or Perplexity. If every assistant underperforms, the prompt is the constraint. If only ChatGPT does on your particular workload, switching or splitting tasks across tools is reasonable, but do it on evidence rather than a general sense of decline.