The AI DownsideDocumenting AI's downsides

AI Slop

Model Collapse: What Happens When AI Trains on AI

More than a third of new web pages already show signs of AI authorship, Pew finds. That is exactly the diet a peer-reviewed study warns can make the next generation of models quietly worse — and blander.

Editorial illustration for “Model Collapse: What Happens When AI Trains on AI”.

Here is a slightly uncomfortable thought experiment. Photocopy a photograph. Now photocopy the photocopy, and the copy of that, and keep going. Within a dozen generations the image is a grey smear: the fine detail goes first, then the midtones, until all that’s left is a muddy average of what was once there. Generative AI has a version of this problem, it has a name, and thanks to a study published in August 2026, we now have a decent measure of how quickly the conditions for it are arriving.

Answer first. Model collapse is what happens when you train a generative model largely on the output of other generative models instead of on human-made data. The model doesn’t explode; it narrows. The rare stuff — unusual facts, minority styles, the long tail of human weirdness — fades out first, and the model drifts toward a confident, bland average. It has been demonstrated in a peer-reviewed Nature paper. And the reason it’s worth your attention now, rather than as a footnote for researchers, is that the web these models learn from is filling up with AI writing: the Pew Research Center reported on 20 August 2026 that more than a third of web pages published since ChatGPT’s launch show signs of AI authorship. The photocopier, in other words, has started feeding on its own copies.

We’re going to be careful here, because this is a topic that attracts both hand-waving doom and lazy dismissal, and neither is honest. Collapse is real and demonstrated; it is also avoidable, and today’s models are not visibly falling apart. The interesting question is the one in between.

The mechanism, without the hand-waving

The core idea is easier than the jargon. Any model trained on data learns an approximation of the distribution that produced it — what’s common, what’s rare, how things vary. When the model then generates new data, it does so slightly imperfectly: it over-represents the likely, under-represents the unlikely, and smooths off the extremes. Feed that output back in as training data for the next model, and the next inherits and amplifies the distortion. Repeat, and you get two failure stages the researchers describe precisely: an early collapse, where the model starts losing information about the rare tails of the distribution, and a late collapse, where those tails vanish entirely and the model converges toward a narrow point that bears less and less resemblance to the original variety.

The definitive account is Ilia Shumailov and colleagues’ 2024 paper in Nature, bluntly titled “AI models collapse when trained on recursively generated data.” They showed the effect across a range of model types, and their conclusion is quotable and stark: the “indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear.” The word doing the work is irreversible. Once a generation of models has been trained past a certain point on its own exhaust, you can’t easily recover the diversity you lost; the rare knowledge isn’t hiding, it’s gone.

One nuance is worth keeping, because it’s where lazy summaries go wrong: collapse in the wild is a matter of degree, not an on/off switch. The dramatic, model-eats-itself version comes from experiments that deliberately loop a model’s output straight back in with nothing else mixed in. Real training pipelines are messier and better defended than that. But the same work shows the pressure operates on a gradient — the more synthetic, unfiltered data creeps into the mix, the more the tails erode — so the useful question isn’t “have the models collapsed?” but “how much variety are we quietly trading away, and can we still tell?”

Why 2026 is when this stopped being theoretical

For a couple of years, model collapse was a tidy lab result with a comforting caveat: nobody actually trains a model exclusively on its own output, so the catastrophic version doesn’t happen in the wild. That caveat is now under pressure, and the Pew study is why.

Pew’s Data Labs ran nearly half a million English-language web pages through an AI-detection model and looked at how the share of likely-AI text has moved over time. The headline figures: about one in ten of all pages sampled in July 2026 showed significant signs of AI authorship, and among pages published since ChatGPT arrived in late 2022, more than a third did. The split by domain is telling — roughly one in ten .com pages showed the signs, against around one in a hundred for .edu and .gov — which is to say the commercial web, the part that gets scraped most eagerly, is precisely the part filling with synthetic text.

The models that write the web are increasingly training on a web written by models. That is not a metaphor for model collapse; it is the exact input condition the research warns about.

A necessary caveat, because we hold ourselves to the sourcing bar we demand of others: “signs of AI authorship” is a probabilistic detector’s judgement, not a confession, and AI-text detection is imperfect and prone to false positives. Pew is measuring likelihood, not counting confessions. But even discounted for that uncertainty, the direction is not seriously in doubt, and there’s a delicious irony in how the detector spots the machines: among the markers that rose most since 2023 were a doubling in em dashes, a jump in a certain kind of vocabulary, and a near-tripling of “negative parallelism” — the “it’s not just X, it’s Y” construction. We use em dashes too; we’ll take the accusation on the chin. The point stands: the fingerprints of generated text are now common enough on the open web to be measured at scale.

The ‘data wall’ makes the temptation worse

There’s an economic engine underneath all this that turns a lab curiosity into a live risk. The frontier labs have spent years training on more or less the entire stock of high-quality, human-written, publicly available text, and there is broad agreement in the field that fresh human writing is now scarce relative to how fast models want to grow — the so-called “data wall.” When you have nearly exhausted the human data, the cheapest and most abundant source left is text that models themselves produced. That is precisely the fuel the collapse research warns against, arriving exactly at the moment the incentive to use it is strongest.

None of which makes synthetic data automatically bad. Labs increasingly generate it on purpose — to cover gaps, balance rare cases and drill specific skills — and done deliberately it helps. The danger is the passive version: hoovering up whatever is cheap and plentiful, which now includes an ever-larger slice of an open web that was itself machine-written. The Pew figure is the measure of how much that slice has grown; the data wall is the reason the industry is tempted to lean on it regardless.

What collapse would actually feel like to you

Forget the sci-fi framing. Model collapse would not announce itself with a crash or a wrong answer you could catch. It would feel like flattening. The same handful of confident, average-sounding answers appearing everywhere. Rarer facts getting harder to surface. Minority dialects, regional knowledge and niche expertise thinning out of the models’ range because they were thin in the training data and got thinner with each recursive pass. Genuine surprise — the unusual-but-correct answer — becoming rarer than the plausible-but-generic one.

That’s corrosive precisely because it’s invisible. A hallucination is at least detectable in principle; you can check it. A slow contraction of what the model even considers possible is not something an individual user can see from a single answer. It compounds problems we’ve written about before: it makes hallucinations harder to notice because the average answer sounds ever more fluent, and it accelerates the dynamic behind AI search making the open web worse, where synthetic summaries crowd out the human pages they were built from. The web gets more text and less signal, and the models trained on it inherit the deficit.

There’s a second-order version that’s worse than personal annoyance. If the models that summarise the web get blander, and the web they summarise is increasingly their own blander output, the loop tightens: fewer readers click through to the messy, specific, human primary source, so fewer such sources are made or seen, so the next model has even less of them to learn from. We’ve traced this from the traffic side in the slow starvation of the open web by AI overviews; model collapse is the same story told from inside the training set, and the two feed each other.

The steel-man: this is a choice, not a doom

Now the fair part, and it’s substantial, because the reverse-hype version of this story — “AI is about to eat itself and collapse” — is just as dishonest as the boosterism we usually push back on. Model collapse is a demonstrated risk, not an inevitability, and the same research that named it also points at the fix.

First, no serious lab trains indiscriminately on raw scraped output. They filter, deduplicate, weight toward trusted sources, and mix in licensed human data and human feedback. Second, and more importantly, the maths of collapse depends on replacing human data with synthetic data; if you keep accumulating fresh human data alongside the synthetic, the collapse is arrested. The Nature authors themselves frame the cure as periodically injecting real human-generated content. Synthetic data isn’t poison in small, curated, well-labelled doses — it’s genuinely useful for filling gaps and augmenting rare cases. The danger is specifically the lazy, high-volume, unprovenanced version: training on whatever is cheap and abundant, which increasingly means other models’ output. Curation is the whole game: a filtered, labelled, minority share of synthetic data is a useful tool, while an unfiltered, unlabelled majority is a slow poison, and the gap between them is just how much effort a lab is willing to spend.

So the honest framing isn’t “the models are doomed.” It’s that avoiding collapse costs money and discipline — paying for human data, tracking provenance, filtering synthetic inputs — and the cheap path runs straight into the rot. This connects to a fight we’ve covered from the other end: the same industry that would benefit from clean, well-labelled human writing is the one arguing hardest that it shouldn’t have to pay for the human writing it already used, as in the unresolved question of who owns the words that trained your AI. A world where human writing is treated as a free, infinite resource is exactly the world where the well gets polluted.

What to watch, and what it means for you

You can’t audit a training set from your armchair, but you can watch for the signals and adjust how much you trust the machine:

  • Notice sameness. If every model gives you the same middle-of-the-road answer in the same cadence, that’s the flattening, not a consensus of experts. Treat convergence as a prompt to go find a primary human source, not as confirmation.
  • Value the tails deliberately. For anything rare, regional, specialist or recent, assume the model is weakest exactly where the training data was thinnest, and verify against a human authority.
  • Reward provenance. Tools and publishers that cite real, human, primary sources are doing the expensive thing that keeps the ecosystem healthy. That is worth paying for, and worth preferring.
  • Keep making human things. The uncomfortable flip side of the Pew number is that human-written, human-verified work is becoming scarcer and therefore more valuable — to readers and, ironically, to the models themselves.

Model collapse is not the end of AI, and anyone selling it as an imminent apocalypse is running the doom version of the same hype machine we distrust. But it is a real, measurable pressure, and August 2026 was the month the input condition — a web increasingly written by machines — stopped being a projection and started being a Pew statistic. The technology that promised to help us make sense of the world’s information has a quiet incentive to fill that world with its own average-sounding output and then learn from it. Whether it avoids drinking from its own well is not a question of physics. It’s a question of what the industry is willing to pay for — and, as ever, of whether anyone makes it.

Frequently asked questions

What exactly is model collapse?

It is the degradation that happens when a generative model learns from data generated by other models rather than from humans. Because each generation slightly over-represents what's common and under-represents what's rare, repeating the loop makes the errors compound: first the 'tails' of the distribution (unusual events, minority styles, uncommon facts) thin out, then they vanish, and the model converges on a narrow, average-sounding core. The 2024 Nature study that named the effect showed it across several model families, from simple statistical models to language models.

Is model collapse actually happening in the real models I use?

The clean, catastrophic version has mostly been shown in controlled experiments where researchers deliberately feed a model its own output over and over. Real labs don't do that on purpose; they filter data, mix in human text and licensed sources, and curate. What has changed in 2026 is the environment: Pew's finding that a large and growing share of new web pages show signs of AI authorship means the open web those models scrape is no longer a clean human sample. So the pressure that causes collapse is now real and rising, even if today's models aren't visibly collapsing.

Why should I care as an ordinary user rather than an AI researcher?

Because the symptom you'd feel isn't an error message, it's blandness and sameness. Model collapse erodes exactly the things that make an answer useful: the rare fact, the unusual perspective, the minority dialect, the long-tail expertise. If the models that write the web, and the models that learn from that web, both drift toward a confident average, you get more text that sounds authoritative and says less, and it becomes harder to find the genuinely surprising or correct-but-uncommon answer. It also compounds problems we've covered before, from hallucinations to the hollowing-out of search.

Can the industry stop it?

Yes, in principle, and this is the fair part. The same research that identified collapse also identified the cure: keep injecting fresh human data rather than replacing it, filter and label synthetic content, and treat provenance as valuable. Some labs pay for licensed human data and human feedback precisely for this reason. The open question is economic, not scientific: clean, human, well-provenanced data is expensive, and the temptation is to train on whatever is cheap and abundant. Whether the industry pays for quality or lets the commons degrade is a choice, not a law of nature.

Sources

  1. AI models collapse when trained on recursively generated data (Shumailov et al., Nature 631, 755–759, 2024)Nature
  2. Author Correction: AI models collapse when trained on recursively generated data (2025)Nature
  3. How Much of the Internet Is Written With AI? (20 August 2026)Pew Research Center
  4. A third of webpages published since ChatGPT's launch show signs of AI authorship, study findsTechCrunch

Related grievances

All articles →