Your favorite AI might be getting dumber, and you are not imagining it. In 2026, people asking “is AI getting dumber?” want to know why ChatGPT or Claude is getting dumber, or simply worse, after a change in a task they depend on.
This article helps you find the practical cause before you change a workflow or pay for a different plan. It separates a real capability change from a model switch, shrinking allowance, overloaded service, long-chat context failure, or a test that did not control its settings.
It traces recent ChatGPT and Claude changes, the 2025 brain rot experiment, and a repeatable diagnostic, and it keeps the product question separate from whether AI is making us dumber, which our brain study covers.
Executive Summary
Quick answer: AI is not proven dumber everywhere. ChatGPT and Claude can feel worse when routing, context, safety settings, usage limits, or outages change. Whether people ask if AI is becoming dumber, getting stupider, or making people stupid, the product and human questions differ. Test a fresh chat before calling a bad week regression.
- Frontier benchmarks are improving: OpenAI reports 53.6 on Agents’ Last Exam and 92.2% on BrowseComp for GPT-5.6, while Anthropic reports 73.6 ± 2.1 for Opus 5 on its Conceptual Reasoning Index (OpenAI results, Anthropic index).
- A Microsoft study across more than 200,000 simulated conversations found a 39% average multi-turn performance drop.
- Product routing, context summaries, safety layers, allowances, and temporary incidents can all make the same app feel less capable. A listed r/ChatGPT thread is an anecdotal report, not proof of a cause.
- Claude status recorded three consecutive days of degradation notices, August 17 to 19, 2026.
- The brain rot result is real controlled research, not proof that current ChatGPT or Claude is decaying.
- Elephas is a privacy-friendly AI knowledge assistant with a Free plan. Elephas starts at $19/month. Try Elephas for free.
Key research findings:
The brain rot study found that clean instruction data worth 4.8x the junk-token count still left a 17.3-point reasoning gap. Recovery was partial, not complete.
- The brain rot study continually pretrained four open models on selected low-quality X/Twitter data. It did not test production ChatGPT or Claude.
- ARC-Challenge with Chain-of-Thought fell from 72.1 to 57.2, while RULER-CWE fell from 83.7 to 52.3, as the junk ratio rose from 0% to 100%.
- The multi-turn study tested 15 models over six tasks and more than 200,000 simulated conversations.
- Documented product limits and incidents should be kept separate from anonymous reports of disappointing answers.
What Is Happening to AI Models?
AI models can get worse on a measured task after training or post-training changes, but a deployed model does not rot merely because people keep using it. The question is, “Is AI itself getting dumber?” or did a consumer product change beneath the same familiar app name?
The multi-turn study found a 39% average drop when complete requests were split across normal conversations. A post in r/ChatGPT called the product “much worse at following instructions for more than 2–3 messages.” That experience fits the study’s context problem, but does not identify a changed model.
Key findings from the research:
- A pinned API snapshot can stay fixed while a consumer default changes.
- A routed fallback can answer differently from the model a user expected.
- Long chats can accumulate assumptions or summarized instructions.
- A poor result can come from product settings without a model-weight change.

Posted in r/Bard in 2025, with 29 upvotes and 55 comments. Usernames are blurred.
Why does ChatGPT feel dumber in 2026?
ChatGPT can feel dumber in 2026 because eight layers affect an answer, and only one is model weights. Has AI gotten dumber, has ChatGPT gotten dumber, or did a product layer move?
OpenAI documents plan-specific models, automatic reasoning, and fallbacks after allowances. Anthropic documents context management, effort controls, and model-specific usage limits (OpenAI help, Claude limits).
| Layer | What can change | What it can look like |
|---|---|---|
| Model weights | A snapshot is replaced or updated | Different knowledge, reasoning, or style |
| Inference configuration | Reasoning effort or context allocation | Missed steps or shorter answers |
| Product routing | Automatic selection or fallback | Same brand, different model |
| Context management | Earlier turns are summarized | Lost instructions or memory |
| Safety layers | A policy classifier or alternate route | Cautious or constrained answers |
| Plan limits | Model access or allowance changes | Preferred capability disappears |
| Transient outages | Errors, latency, or failed tools | A bad day that later resolves |
| Hidden quantization or throttling | Undisclosed precision or compute changes | Unverified for consumer ChatGPT and Claude |
- Check the visible model and plan before blaming the weights.
- Start a new chat before judging a long conversation.
- Fix reasoning effort and tools for a fair comparison.
- Check a vendor status page before calling a bad result permanent.

Posted in r/ChatGPT in 2025, with 48 upvotes and 23 comments. Usernames are blurred.
A subscriber in r/ChatGPTcomplaints reported the system was “suddenly forgetting things and being less competent.” Anthropic documents that older messages can be summarized as a conversation approaches context limits, which is a documented explanation for some memory complaints.
What changed in ChatGPT and Claude in the last 90 days?
Thirteen documented changes between May 28 and August 14, 2026 can affect perceived quality without proving a field-wide decline. Each row comes from a first-party source. The table separates releases, routing, and entitlement changes from theories that vendors have not confirmed. A side-by-side of the major assistants is a better basis than one bad chat.
| Date | Product | Change | Why it can feel worse | Source |
|---|---|---|---|---|
| May 28 | ChatGPT | Default style updated; retirements announced | Style can change usefulness | Release notes |
| May 28 | Claude | Opus 4.8 launched | A model swap changes behavior | Opus launch |
| June 9 | Claude | Fable 5 and Mythos 5 launched with safety routing | Some prompts can follow another route | Fable launch |
| June 24 | ChatGPT | GPT-5.5 Instant replaced GPT-5.3 Instant | Existing chats can behave differently | Release notes |
| June 30 | Claude | Sonnet 5 became the Free and Pro default | Default answers come from another model | Sonnet launch |
| June 30 | Claude | Fable 5 was redeployed | One visible name can change behavior | Redeployment |
| July 6 | ChatGPT | GPT-5.5 Instant Mini became a fallback | Lower-tier fallback after allowance use | Release notes |
| July 9 | ChatGPT | GPT-5.6 Sol launched with product changes | Model, tools, and interface changed together | OpenAI launch |
| July 20 | Claude | Fable left Pro’s included limits for credits | Premium access shrank without a weight change | Fable rules |
| July 24 | Claude | Opus 5 added effort levels | Effort became a quality control | Platform notes |
| July 25 to 27 | ChatGPT | Elevated errors were reported, then resolved | An incident is not long-term decline | OpenAI status |
| August 6 | ChatGPT | Sol reached eligible paid plans; Luna became Free/Go default | One brand maps to different models by plan | OpenAI help |
| August 14 | ChatGPT | Think and project-memory controls expanded | Reasoning and memory became visible controls | Release notes |
- The August 6 Sol and Luna split can explain plan-specific reports.
- The July 6 fallback is documented, not a secret downgrade.
- The July 20 Fable change is entitlement shrinkflation, not proof of worse weights.
- The July 25 to 27 incident was resolved and should not be treated as model collapse.
On July 20, Fable moved from Claude Pro’s included usage to credits. Max users could use Fable for up to 50% of their shared weekly limit (Fable rules). A listed r/Anthropic discussion shows why readers should track plan changes separately from model-quality claims.
The next section separates Claude’s three-day degradation sequence, August 17 to 19, from the August 16 service disruption. Its Claude status record belongs in the transient-outage layer, not a claim about brain rot or model weights.
Is Claude getting dumber too?
Claude can feel worse for the same practical reasons as ChatGPT: a changed default, context summary, effort setting, safety route, allowance, or incident. From August 17 through 19, 2026, Anthropic posted three consecutive days of degradation notices. August 16 was a separate service disruption, not a model-degradation notice.
- Sonnet 5 became the default for Free and Pro on June 30.
- Claude usage is shared across its web, desktop, and coding products.
- Earlier turns can be summarized as context fills.
- Pinned model identifiers are better for repeatable tests than a broad product label.
- The August 17 to 19 notices are a transient-outage signal, not a weight-change finding.
Since August 1, the status history listed 7 incidents titled “Degraded performance” and 12 other notices, including Service disruption, Elevated errors, Error rates, and an issue reaching the status page. The August 16 critical disruption affected claude.ai, platform.claude.com, the Claude API, Claude Code, and Claude Cowork (Claude status).
| Date (UTC) | Impact | Incident title | Window | Source |
|---|---|---|---|---|
| Aug 17 | minor | Degraded performance for Claude Opus 5, Claude Sonnet 5 | 13:56 to 15:29 | Claude status |
| Aug 18 | minor | Degraded performance for multiple models | 16:11 to 18:23 | Claude status |
| Aug 19 | minor | Degraded performance for Claude Opus 5 and Claude Haiku 4.5 | 09:42 to 11:02 | Claude status |

Anthropic has published no root cause for this group of incidents. Any capacity or compute explanation is reader speculation and remains unverified. Anthropic publishes incidents and resolves them in hours, an important transparency point, but the notices do not show what caused them.
- r/ClaudeCode’s “Anthropic has nerfed every model” thread had 200+ comments.
- r/ClaudeAI’s “Opus 5 is actually almost rage-inducing to use.” thread had 440+ comments.
- r/ClaudeAI hosted Discussion Hub threads for the Aug 16 incident, Aug 17 incident, and Aug 18 incident. The Aug 16 hub had 50+ comments and the Aug 18 hub had 100+.
Threads answer the “is it just me or is AI getting dumber?” question with similar accounts, but perceived model quality is a claim Anthropic has not confirmed. One r/Anthropic post said, “Opus 4.8 Max now often feels worse than old Haiku.” That is an anecdote, not an audit.
Understanding Brain Rot in Simple Terms
AI brain rot is the nickname for a controlled training-data effect. People use the term for shallow online content that weakens attention. In the study, it means that repeatedly training an LLM on selected junk text changed its reasoning and long-context performance. It does not mean your individual chats retrain a public model.
The researchers built junk and control sets from X/Twitter using engagement and semantic-quality measures. Their experiment tested Llama and Qwen models, not consumer ChatGPT or Claude. A free-tier r/ChatGPT post said it “feels like the free model was downgraded,” but that report cannot establish the study’s training mechanism.
How brain rot affects AI differently:
- Human brain rot is a cultural phrase about focus and attention.
- AI brain rot refers to a continual-training intervention.
- The observed effect involved reasoning, long context, safety, and personality measures.
- The study is a warning about data quality, not proof of current closed-model decay.
How the Scientists Tested This Theory
The scientists tested whether selected low-quality training data could cause measurable decline under controlled conditions. They continually pretrained four open models on matched X/Twitter data mixtures, then evaluated reasoning, long-context understanding, safety, and personality. Production ChatGPT and Claude were not study subjects.

The current arXiv version reports ARC-Challenge Chain-of-Thought falling from 72.1 to 57.2 and RULER-CWE falling from 83.7 to 52.3 (study abstract). Version one, published in October 2025, reported 74.9 and 84.4, which is why many 2025 stories and the retained results image use those earlier figures.
Across the arXiv study, declines increased with the junk share: at 20%, ARC-C CoT moved about 0.9 points and RULER-CWE about 0.4 points from baseline, while the steepest declines came at 80% to 100%.
The experimental setup included:
- Four primary open models, plus a Llama 3 70B follow-up.
- Training mixtures ranging from 0% to 100% selected junk content.
- Matched token scale and training operations across conditions.
- Instruction tuning after continual pretraining.

A document-review complaint in r/ChatGPTcomplaints asked, “Final QA this document, no recommendations.” A single answer that ignores that instruction is not evidence of brain rot. It does show why instruction-following should be tested with a fixed prompt.
The Biggest Problem: Skipping Steps in Thinking
The study identified thought skipping as the main reasoning failure. As the junk-data ratio rose, models increasingly shortened or skipped the intermediate reasoning needed to solve a problem. A model can hold relevant facts and still fail because it jumps to a conclusion before using them well.
Why thought skipping matters:
- It accounted for the majority of reasoning errors in affected models.
- Models trained on clean data consistently showed complete reasoning chains.
- The skipping got worse as junk data percentage increased.
- Even when models had the right knowledge, they failed by not using proper logic.
Models cut reasoning short or skipped whole steps, and the change was in how they approached problems, not only which answers they got. They could hold relevant facts but jump to a conclusion before using them. This failure mode explained much of the performance drop across the tested tasks.
Thought skipping also helps explain why a chat can feel inconsistent. A rushed or low-effort response may omit checks that a longer response would perform. That is why a test should fix the model, effort setting, tools, prompt, and conversation state before comparing outputs.
An r/ChatGPT thread about poor instruction following in longer chats and an r/ChatGPTcomplaints report are symptom reports, not evidence of training-data damage. The latter said, “Final QA this document, no recommendations.” Both support controlled tests for skipped steps, context loss, and changed settings before assigning a cause.
Why This Problem Doesn't Go Away Easily
The researchers tried to repair affected models with clean instruction data and further instruction tuning. Recovery was partial. The study found that even after instruction data equal to 4.8x the junk-training tokens, the best mitigated models still had a 17.3-point ARC-C Chain-of-Thought gap and a 9.0-point RULER gap (paper).
Why recovery falls short:
- The junk data causes something called representational drift in the model's internal structure.
- This drift changes how the model fundamentally understands and processes information.
- Retraining addresses surface-level patterns but doesn't fix deeper structural changes.
- Mitigation was tested against the 100% junk condition and still left a measurable gap.
The finding matters for people training or selecting models, not as a verdict on every disappointing chat. A listed r/ChatGPTcomplaints discussion may point to a real workflow problem, but it cannot reveal a provider’s private training data.
Is AI training on its own output?
Model collapse is a related training-data risk, but it is not proof that ChatGPT or Claude is collapsing right now. Model collapse occurs when generated data recursively replaces the original real data across generations. It is a specific failure condition, not an automatic outcome whenever a model uses some synthetic data.
- Recursive replacement is the central risk.
- Keeping original real data changed the result in later experiments.
- Adequate verification of synthetic data can mitigate or slow collapse under the study conditions.
- No public source establishes current collapse in a named closed consumer model.
The Nature study showed information loss under recursive replacement. Later work on retained data and synthetic verification found ways to avoid or slow that outcome. Does AI get dumber over time? It can under certain training choices, but inevitability is not supported.
How can you tell if ChatGPT actually got worse?
A reproducible quality test compares the same task under controlled conditions. Start with a new chat, pin the model where possible, choose one fixed effort setting, turn tools off, and run several trials. Save the transcript, plan, remaining allowance, and time before drawing a conclusion.
- Use a small test set of real but de-identified tasks.
- Keep the prompt, expected output, and scoring method fixed.
- Record the visible model, plan, tools, effort, and allowance.
- Check the vendor status page before and after a bad result.
- Compare a fresh chat with the long conversation that felt worse.
Claude Code moderators opened a model-performance megathread to collect model details and reproducible prompts. That is the useful standard. A saved test can separate a genuine regression from context buildup, a fallback, or an outage.
How to Protect Your AI From Brain Rot
You cannot control a cloud vendor’s training mixture, routing, or allowance rules. A listed r/ChatGPTcomplaints discussion is an anecdotal report, not evidence of a training cause. You can control what information an AI uses for important work.
Elephas is a private AI knowledge assistant for Mac that redacts sensitive data before it reaches cloud models. Super Brain is the knowledge base built from your documents; Super Chat uses it for cited answers. Elephas provides built-in local LLM models for offline Mac work.
For people handling sensitive documents who still want a cloud model, Elephas adds automatic PII redaction. Before a prompt is sent to ChatGPT, Claude, Gemini, Grok, Perplexity, or any other cloud model, Elephas strips sensitive names, emails, phone numbers, and identifiers on your Mac.
The cloud model only ever sees the sanitized text. When the answer comes back, the redacted fields are reassembled locally on your machine, so identifiable information never leaves the device. Elephas pairs this with zero data retention: content never trains AI models, never sits on a vendor's server, and never passes through a third-party reviewer's screen.

The practical solution:
- Build a curated knowledge base from documents you trust.
- Use cited answers to check where an important claim came from.
- Keep confidential source material local when possible.
- Smart Redaction is available on every plan, including Free.
- Elephas has a Free plan and starts at $19/month. Try Elephas for free.
Elephas does not freeze a cloud model or make every answer correct. It gives you a stable source base, a record to verify, and a privacy boundary for sensitive work. That is more useful than treating anonymous reports of worse answers as proof that every AI assistant has become stupid.
