Is AI Getting Dumber in 2026? Why ChatGPT and Claude Can Feel Worse

Your favorite AI might be getting dumber, and you are not imagining it. In 2026, people asking “is AI getting dumber?” want to know why ChatGPT or Claude is getting dumber, or simply worse, after a change in a task they depend on.

This article helps you find the practical cause before you change a workflow or pay for a different plan. It separates a real capability change from a model switch, shrinking allowance, overloaded service, long-chat context failure, or a test that did not control its settings.

It traces recent ChatGPT and Claude changes, the 2025 brain rot experiment, and a repeatable diagnostic, and it keeps the product question separate from whether AI is making us dumber, which our brain study covers.

Executive Summary

Quick answer: AI is not proven dumber everywhere. ChatGPT and Claude can feel worse when routing, context, safety settings, usage limits, or outages change. Whether people ask if AI is becoming dumber, getting stupider, or making people stupid, the product and human questions differ. Test a fresh chat before calling a bad week regression.

  • Frontier benchmarks are improving: OpenAI reports 53.6 on Agents’ Last Exam and 92.2% on BrowseComp for GPT-5.6, while Anthropic reports 73.6 ± 2.1 for Opus 5 on its Conceptual Reasoning Index (OpenAI results, Anthropic index).
  • A Microsoft study across more than 200,000 simulated conversations found a 39% average multi-turn performance drop.
  • Product routing, context summaries, safety layers, allowances, and temporary incidents can all make the same app feel less capable. A listed r/ChatGPT thread is an anecdotal report, not proof of a cause.
  • Claude status recorded three consecutive days of degradation notices, August 17 to 19, 2026.
  • The brain rot result is real controlled research, not proof that current ChatGPT or Claude is decaying.
  • Elephas is a privacy-friendly AI knowledge assistant with a Free plan. Elephas starts at $19/month. Try Elephas for free.

Key research findings:

The brain rot study found that clean instruction data worth 4.8x the junk-token count still left a 17.3-point reasoning gap. Recovery was partial, not complete.

  • The brain rot study continually pretrained four open models on selected low-quality X/Twitter data. It did not test production ChatGPT or Claude.
  • ARC-Challenge with Chain-of-Thought fell from 72.1 to 57.2, while RULER-CWE fell from 83.7 to 52.3, as the junk ratio rose from 0% to 100%.
  • The multi-turn study tested 15 models over six tasks and more than 200,000 simulated conversations.
  • Documented product limits and incidents should be kept separate from anonymous reports of disappointing answers.

What Is Happening to AI Models?

AI models can get worse on a measured task after training or post-training changes, but a deployed model does not rot merely because people keep using it. The question is, “Is AI itself getting dumber?” or did a consumer product change beneath the same familiar app name?

The multi-turn study found a 39% average drop when complete requests were split across normal conversations. A post in r/ChatGPT called the product “much worse at following instructions for more than 2–3 messages.” That experience fits the study’s context problem, but does not identify a changed model.

Diagram of the LLM brain rot hypothesis and controlled experiment

Key findings from the research:

  • A pinned API snapshot can stay fixed while a consumer default changes.
  • A routed fallback can answer differently from the model a user expected.
  • Long chats can accumulate assumptions or summarized instructions.
  • A poor result can come from product settings without a model-weight change.
Reddit post in r/Bard reporting that Google's AI models forget conversation context and give less accurate answers

Posted in r/Bard in 2025, with 29 upvotes and 55 comments. Usernames are blurred.

Why does ChatGPT feel dumber in 2026?

ChatGPT can feel dumber in 2026 because eight layers affect an answer, and only one is model weights. Has AI gotten dumber, has ChatGPT gotten dumber, or did a product layer move?

OpenAI documents plan-specific models, automatic reasoning, and fallbacks after allowances. Anthropic documents context management, effort controls, and model-specific usage limits (OpenAI help, Claude limits).

LayerWhat can changeWhat it can look like
Model weightsA snapshot is replaced or updatedDifferent knowledge, reasoning, or style
Inference configurationReasoning effort or context allocationMissed steps or shorter answers
Product routingAutomatic selection or fallbackSame brand, different model
Context managementEarlier turns are summarizedLost instructions or memory
Safety layersA policy classifier or alternate routeCautious or constrained answers
Plan limitsModel access or allowance changesPreferred capability disappears
Transient outagesErrors, latency, or failed toolsA bad day that later resolves
Hidden quantization or throttlingUndisclosed precision or compute changesUnverified for consumer ChatGPT and Claude
  • Check the visible model and plan before blaming the weights.
  • Start a new chat before judging a long conversation.
  • Fix reasoning effort and tools for a fair comparison.
  • Check a vendor status page before calling a bad result permanent.
Reddit post in r/ChatGPT asking whether OpenAI has addressed a recent drop in ChatGPT answer quality

Posted in r/ChatGPT in 2025, with 48 upvotes and 23 comments. Usernames are blurred.

A subscriber in r/ChatGPTcomplaints reported the system was “suddenly forgetting things and being less competent.” Anthropic documents that older messages can be summarized as a conversation approaches context limits, which is a documented explanation for some memory complaints.

What changed in ChatGPT and Claude in the last 90 days?

Thirteen documented changes between May 28 and August 14, 2026 can affect perceived quality without proving a field-wide decline. Each row comes from a first-party source. The table separates releases, routing, and entitlement changes from theories that vendors have not confirmed. A side-by-side of the major assistants is a better basis than one bad chat.

DateProductChangeWhy it can feel worseSource
May 28ChatGPTDefault style updated; retirements announcedStyle can change usefulnessRelease notes
May 28ClaudeOpus 4.8 launchedA model swap changes behaviorOpus launch
June 9ClaudeFable 5 and Mythos 5 launched with safety routingSome prompts can follow another routeFable launch
June 24ChatGPTGPT-5.5 Instant replaced GPT-5.3 InstantExisting chats can behave differentlyRelease notes
June 30ClaudeSonnet 5 became the Free and Pro defaultDefault answers come from another modelSonnet launch
June 30ClaudeFable 5 was redeployedOne visible name can change behaviorRedeployment
July 6ChatGPTGPT-5.5 Instant Mini became a fallbackLower-tier fallback after allowance useRelease notes
July 9ChatGPTGPT-5.6 Sol launched with product changesModel, tools, and interface changed togetherOpenAI launch
July 20ClaudeFable left Pro’s included limits for creditsPremium access shrank without a weight changeFable rules
July 24ClaudeOpus 5 added effort levelsEffort became a quality controlPlatform notes
July 25 to 27ChatGPTElevated errors were reported, then resolvedAn incident is not long-term declineOpenAI status
August 6ChatGPTSol reached eligible paid plans; Luna became Free/Go defaultOne brand maps to different models by planOpenAI help
August 14ChatGPTThink and project-memory controls expandedReasoning and memory became visible controlsRelease notes
  • The August 6 Sol and Luna split can explain plan-specific reports.
  • The July 6 fallback is documented, not a secret downgrade.
  • The July 20 Fable change is entitlement shrinkflation, not proof of worse weights.
  • The July 25 to 27 incident was resolved and should not be treated as model collapse.

On July 20, Fable moved from Claude Pro’s included usage to credits. Max users could use Fable for up to 50% of their shared weekly limit (Fable rules). A listed r/Anthropic discussion shows why readers should track plan changes separately from model-quality claims.

The next section separates Claude’s three-day degradation sequence, August 17 to 19, from the August 16 service disruption. Its Claude status record belongs in the transient-outage layer, not a claim about brain rot or model weights.

Is Claude getting dumber too?

Claude can feel worse for the same practical reasons as ChatGPT: a changed default, context summary, effort setting, safety route, allowance, or incident. From August 17 through 19, 2026, Anthropic posted three consecutive days of degradation notices. August 16 was a separate service disruption, not a model-degradation notice.

  • Sonnet 5 became the default for Free and Pro on June 30.
  • Claude usage is shared across its web, desktop, and coding products.
  • Earlier turns can be summarized as context fills.
  • Pinned model identifiers are better for repeatable tests than a broad product label.
  • The August 17 to 19 notices are a transient-outage signal, not a weight-change finding.

Since August 1, the status history listed 7 incidents titled “Degraded performance” and 12 other notices, including Service disruption, Elevated errors, Error rates, and an issue reaching the status page. The August 16 critical disruption affected claude.ai, platform.claude.com, the Claude API, Claude Code, and Claude Cowork (Claude status).

Date (UTC)ImpactIncident titleWindowSource
Aug 17minorDegraded performance for Claude Opus 5, Claude Sonnet 513:56 to 15:29Claude status
Aug 18minorDegraded performance for multiple models16:11 to 18:23Claude status
Aug 19minorDegraded performance for Claude Opus 5 and Claude Haiku 4.509:42 to 11:02Claude status
Claude Status incident history for August 2026, showing degraded performance notices for Claude Opus 5, Sonnet 5 and Haiku 4.5 on August 17, 18 and 19, followed by separate service disruption and elevated error incidents

Anthropic has published no root cause for this group of incidents. Any capacity or compute explanation is reader speculation and remains unverified. Anthropic publishes incidents and resolves them in hours, an important transparency point, but the notices do not show what caused them.

  • r/ClaudeCode’s “Anthropic has nerfed every model” thread had 200+ comments.
  • r/ClaudeAI’s “Opus 5 is actually almost rage-inducing to use.” thread had 440+ comments.
  • r/ClaudeAI hosted Discussion Hub threads for the Aug 16 incident, Aug 17 incident, and Aug 18 incident. The Aug 16 hub had 50+ comments and the Aug 18 hub had 100+.

Threads answer the “is it just me or is AI getting dumber?” question with similar accounts, but perceived model quality is a claim Anthropic has not confirmed. One r/Anthropic post said, “Opus 4.8 Max now often feels worse than old Haiku.” That is an anecdote, not an audit.

Understanding Brain Rot in Simple Terms

AI brain rot is the nickname for a controlled training-data effect. People use the term for shallow online content that weakens attention. In the study, it means that repeatedly training an LLM on selected junk text changed its reasoning and long-context performance. It does not mean your individual chats retrain a public model.

The researchers built junk and control sets from X/Twitter using engagement and semantic-quality measures. Their experiment tested Llama and Qwen models, not consumer ChatGPT or Claude. A free-tier r/ChatGPT post said it “feels like the free model was downgraded,” but that report cannot establish the study’s training mechanism.

Illustration of low-quality content feeding into an AI model

How brain rot affects AI differently:

  • Human brain rot is a cultural phrase about focus and attention.
  • AI brain rot refers to a continual-training intervention.
  • The observed effect involved reasoning, long context, safety, and personality measures.
  • The study is a warning about data quality, not proof of current closed-model decay.

How the Scientists Tested This Theory

The scientists tested whether selected low-quality training data could cause measurable decline under controlled conditions. They continually pretrained four open models on matched X/Twitter data mixtures, then evaluated reasoning, long-context understanding, safety, and personality. Production ChatGPT and Claude were not study subjects.

Effect sizes across reasoning, long-context, safety and personality measures
Table of cognitive functions and the benchmarks used to test each

The current arXiv version reports ARC-Challenge Chain-of-Thought falling from 72.1 to 57.2 and RULER-CWE falling from 83.7 to 52.3 (study abstract). Version one, published in October 2025, reported 74.9 and 84.4, which is why many 2025 stories and the retained results image use those earlier figures.

Across the arXiv study, declines increased with the junk share: at 20%, ARC-C CoT moved about 0.9 points and RULER-CWE about 0.4 points from baseline, while the steepest declines came at 80% to 100%.

The experimental setup included:

  • Four primary open models, plus a Llama 3 70B follow-up.
  • Training mixtures ranging from 0% to 100% selected junk content.
  • Matched token scale and training operations across conditions.
  • Instruction tuning after continual pretraining.
Full results table by junk ratio for reasoning, long-context, safety and personality

A document-review complaint in r/ChatGPTcomplaints asked, “Final QA this document, no recommendations.” A single answer that ignores that instruction is not evidence of brain rot. It does show why instruction-following should be tested with a fixed prompt.

The Biggest Problem: Skipping Steps in Thinking

The study identified thought skipping as the main reasoning failure. As the junk-data ratio rose, models increasingly shortened or skipped the intermediate reasoning needed to solve a problem. A model can hold relevant facts and still fail because it jumps to a conclusion before using them well.

Chart of failure counts showing thought skipping as the dominant failure mode

Why thought skipping matters:

  • It accounted for the majority of reasoning errors in affected models.
  • Models trained on clean data consistently showed complete reasoning chains.
  • The skipping got worse as junk data percentage increased.
  • Even when models had the right knowledge, they failed by not using proper logic.

Models cut reasoning short or skipped whole steps, and the change was in how they approached problems, not only which answers they got. They could hold relevant facts but jump to a conclusion before using them. This failure mode explained much of the performance drop across the tested tasks.

Thought skipping also helps explain why a chat can feel inconsistent. A rushed or low-effort response may omit checks that a longer response would perform. That is why a test should fix the model, effort setting, tools, prompt, and conversation state before comparing outputs.

An r/ChatGPT thread about poor instruction following in longer chats and an r/ChatGPTcomplaints report are symptom reports, not evidence of training-data damage. The latter said, “Final QA this document, no recommendations.” Both support controlled tests for skipped steps, context loss, and changed settings before assigning a cause.

Why This Problem Doesn't Go Away Easily

The researchers tried to repair affected models with clean instruction data and further instruction tuning. Recovery was partial. The study found that even after instruction data equal to 4.8x the junk-training tokens, the best mitigated models still had a 17.3-point ARC-C Chain-of-Thought gap and a 9.0-point RULER gap (paper).

Chart showing brain rot persists against instruction tuning and control data

Why recovery falls short:

  • The junk data causes something called representational drift in the model's internal structure.
  • This drift changes how the model fundamentally understands and processes information.
  • Retraining addresses surface-level patterns but doesn't fix deeper structural changes.
  • Mitigation was tested against the 100% junk condition and still left a measurable gap.

The finding matters for people training or selecting models, not as a verdict on every disappointing chat. A listed r/ChatGPTcomplaints discussion may point to a real workflow problem, but it cannot reveal a provider’s private training data.

Is AI training on its own output?

Model collapse is a related training-data risk, but it is not proof that ChatGPT or Claude is collapsing right now. Model collapse occurs when generated data recursively replaces the original real data across generations. It is a specific failure condition, not an automatic outcome whenever a model uses some synthetic data.

  • Recursive replacement is the central risk.
  • Keeping original real data changed the result in later experiments.
  • Adequate verification of synthetic data can mitigate or slow collapse under the study conditions.
  • No public source establishes current collapse in a named closed consumer model.

The Nature study showed information loss under recursive replacement. Later work on retained data and synthetic verification found ways to avoid or slow that outcome. Does AI get dumber over time? It can under certain training choices, but inevitability is not supported.

How can you tell if ChatGPT actually got worse?

A reproducible quality test compares the same task under controlled conditions. Start with a new chat, pin the model where possible, choose one fixed effort setting, turn tools off, and run several trials. Save the transcript, plan, remaining allowance, and time before drawing a conclusion.

  • Use a small test set of real but de-identified tasks.
  • Keep the prompt, expected output, and scoring method fixed.
  • Record the visible model, plan, tools, effort, and allowance.
  • Check the vendor status page before and after a bad result.
  • Compare a fresh chat with the long conversation that felt worse.

Claude Code moderators opened a model-performance megathread to collect model details and reproducible prompts. That is the useful standard. A saved test can separate a genuine regression from context buildup, a fallback, or an outage.

How to Protect Your AI From Brain Rot

You cannot control a cloud vendor’s training mixture, routing, or allowance rules. A listed r/ChatGPTcomplaints discussion is an anecdotal report, not evidence of a training cause. You can control what information an AI uses for important work.

Elephas is a private AI knowledge assistant for Mac that redacts sensitive data before it reaches cloud models. Super Brain is the knowledge base built from your documents; Super Chat uses it for cited answers. Elephas provides built-in local LLM models for offline Mac work.

For people handling sensitive documents who still want a cloud model, Elephas adds automatic PII redaction. Before a prompt is sent to ChatGPT, Claude, Gemini, Grok, Perplexity, or any other cloud model, Elephas strips sensitive names, emails, phone numbers, and identifiers on your Mac.

The cloud model only ever sees the sanitized text. When the answer comes back, the redacted fields are reassembled locally on your machine, so identifiable information never leaves the device. Elephas pairs this with zero data retention: content never trains AI models, never sits on a vendor's server, and never passes through a third-party reviewer's screen.

Smart Redaction on a Mac: identifiers are stripped before a prompt reaches a cloud model, then reassembled locally
Elephas app with Redact before send switched on, showing the sanitized prompt the cloud model receives

The practical solution:

  • Build a curated knowledge base from documents you trust.
  • Use cited answers to check where an important claim came from.
  • Keep confidential source material local when possible.
  • Smart Redaction is available on every plan, including Free.
  • Elephas has a Free plan and starts at $19/month. Try Elephas for free.

Elephas does not freeze a cloud model or make every answer correct. It gives you a stable source base, a record to verify, and a privacy boundary for sensitive work. That is more useful than treating anonymous reports of worse answers as proof that every AI assistant has become stupid.

Frequently Asked Questions

Is my ChatGPT the same model as everyone else's?
Not necessarily. OpenAI documents different plan access, including Sol for eligible paid plans and Luna for Free and Go users. Automatic routing and post-allowance fallbacks can also change the model path. Record your plan, visible model, and remaining allowance before comparing your result with another person’s.
Does starting a new chat actually fix a bad answer?
A new chat can fix errors caused by accumulated context, prior assumptions, or summarized messages. It cannot repair a changed model or an outage. Run the same prompt in a fresh chat several times, then compare those results with the original conversation before deciding whether the quality changed.
Can a safety rule make a model seem less capable?
Yes. A safety classifier, policy update, or alternate route can make an answer more cautious, shorter, or more limited for a particular request. That may be appropriate for risk management, but it can still feel like a quality drop when compared with an earlier response to a similar prompt.
Would the API give more consistent results than the app?
Often, an API can be easier to test because you can save a model identifier, temperature, prompt, and output limit. It still does not guarantee permanent behavior. Providers may retire snapshots or change product terms, so keep dated transcripts and test results for important workflows.
Does my own chat history retrain the model?
Your chat can affect its own context and remembered instructions, not continually retrain a deployed model. Training controls differ by vendor, so check your account privacy or data settings rather than assume a default. Anthropic says chat training is opt-in. Business, team, and API tiers use terms separate from consumer accounts.
Are benchmark gains proof that a chat will work well?
No. Benchmarks measure a stated model under a stated evaluation setup. Real work can involve long context, vague instructions, tools, plan restrictions, and changing defaults. Use vendor benchmarks as one input, then test your own approved tasks with saved prompts and clear scoring rules.

Sign up now

Get a deep dive into the most important AI story of the week. Deliverd to your inbox for free!

Ayush Chaturvedi

Ayush Chaturvedi, co-founder of Elephas, writes articles on AI to help knowledge workers. He created Elephas, a desktop AI writing assistant for Mac users, to improve productivity and knowledge management. Ayush believes AI can augment human creativity and recommends Elephas Super Brain for personal growth.

Kamban S

Kamban is the founder of Elephas, a native Mac app for seamless AI writing. He writes articles on the latest AI developments and is fueled by his passion for AI's potential. Kamban is committed to user experience and enthusiastic about the future of AI in education and data-driven decision-making. His goal? To make AI user-friendly for everyone.

Jc Chaithanya

Jc Chaithanya
Chaithanya is a freelance content writer passionate about exploring the world of AI and technology. He has a talent for turning complex ideas into clear, engaging content. When not writing, you can find him enjoying the latest anime, drawing inspiration from each episode.

Elephas

Meet Elephas - Your AI-Powered Knowledge Assistant. Your Personal ChatGPT for all your files. Transform information overload into actionable insights. Organize vast knowledge. Access ideas efortlessly. Save 10 hours a week