The Death of the Open Web Through AI Overviews
You cannot indefinitely summarise a web you are simultaneously starving of traffic.
There is a quiet contradiction at the centre of the AI search era, and no amount of clever engineering makes it go away. AI-generated answers are assembled from the open web. The open web is funded, however imperfectly, by people visiting it. AI-generated answers exist to stop people needing to visit it. You do not have to be an economist to notice that this arrangement eats its own supply.
The bargain that used to work
The old search bargain was extractive but survivable. Google crawled your site for free and, in exchange, sent you readers. Those readers saw your ads, bought your things, subscribed to your newsletter, or simply knew you existed for next time. It was not charity, but it was a loop that closed: value in, value out.
Answer engines break the loop. They still take the content — that has not changed — but the return visit is now optional, and increasingly rare, because the answer has already been delivered on the results page. The crawler still arrives. The reader does not.
“Fair use” is doing a great deal of lifting
The legal scaffolding holding this arrangement up is shakier than the confidence of the companies relying on it. The industry's working assumption is that training on publicly available material is permitted — that ingesting the web to build a model is a transformative use that requires no permission and no payment. That assumption is being tested in courtrooms rather than settled in them, and the outcome is genuinely uncertain. A great deal of capital has been deployed on the premise that the answer will be favourable, which is a bold way to build an industry.
What is not in doubt is the asymmetry of the current moment. The party doing the summarising is large, well-funded and organised. The parties being summarised are numerous, scattered and individually powerless — the small publisher has no realistic means of auditing whether their work trained a model, let alone extracting anything for it. Even the licensing deals that have appeared tend to favour the few outlets big enough to command a negotiation, leaving the long tail of the web — the part that actually made the answers good — with the same crawl and none of the cheque. However the law lands, the distribution of leverage tells you who the system was built to serve.
The independent site feels it first
Large publishers with brand recognition and direct traffic can absorb some of this. The people who cannot are exactly the ones the web was best at supporting: the hobbyist who documented an obscure repair, the specialist blogger, the small forum, the person who wrote the definitive guide to one narrow thing because they cared. That long tail of genuine expertise is precisely what makes generated answers good — and it is precisely the part with the thinnest financial cushion.
Strip the traffic from those sites and two things happen. The obvious one: some of them stop, because running a website that nobody visits is a hobby with a hosting bill. The subtler one: the ones that remain start writing for the machine rather than the human, because the machine is now the only reliable reader. Neither outcome improves the thing being summarised.
The feedback loop nobody wants to model out loud
Follow the logic forward and it gets uncomfortable. Answer engines degrade the incentive to publish. Fewer people publish, or publish worse, thinner, more machine-shaped material. The models trained on that thinner web produce thinner answers. And a growing share of what remains online is itself machine-generated — models increasingly learning from the output of models, a copy of a copy of a copy.
You end up with an information ecosystem where the primary sources quietly wither while the summarisation layer on top gets glossier and more confident. It looks healthier than ever from the results page. Underneath, the soil is being mined.
What would actually fix it
The honest fixes are unglamorous and involve the platforms giving something back:
- Attribution that drives clicks, not just citations in small grey text — links positioned and weighted so that being a source is worth something again.
- Real, negotiated compensation for the publishers whose work materially trains and grounds these systems, rather than one-off deals for the largest players only.
- An honest default that shows an answer when the question is settled and gets out of the way when it is not.
The model-collapse problem is not science fiction
There is a technical failure mode lurking at the end of this road, and it has a name: model collapse. As more of the web becomes machine-generated — filler articles, synthetic summaries, content produced to feed the very systems that summarise it — the models trained on that web increasingly learn from their own kind of output rather than from fresh human writing. Train a model on the exhaust of models, repeatedly, and quality degrades in measurable ways: the range narrows, the errors compound, the oddities calcify. A copy of a copy of a copy loses something each generation, and researchers have shown this is not a vague worry but a reproducible effect.
Human-written primary material is the antidote — the fresh signal that keeps the whole apparatus honest. Which makes it darkly ironic that the answer engines depend on exactly the human web they are defunding. The system needs a steady supply of real people writing real things, and it is busy removing the reason those people had to publish. It is eating the seed corn and calling the meal progress.
Who pays, and who decides
Step back and the shape is a familiar one: a platform positioned between producers and readers, extracting value from both while answering to neither. The independent publisher has no seat at the table where the terms are set, no ability to audit whether their work trained a model, and no realistic recourse when the traffic dries up. The reader, meanwhile, is offered convenience and told not to worry about the plumbing. Both are being managed; only the platform is deciding.
That is the part that deserves to be said without hedging. The open web was a genuinely radical thing — anyone could publish, and being useful was rewarded with an audience. Reorganising it so that usefulness is harvested upward and the audience is retained at the top is not a natural evolution. It is a choice, made by a small number of large companies, presented to everyone else as weather. Weather is no one's fault. This has an address.
The reader loses a web worth having
It is easy to frame this as a fight between platforms and publishers and forget the third party with the most to lose: the reader. The open web, for all its spam and clutter, gave people something genuinely valuable — a plurality of voices, the ability to follow a link to a primary source, to read the specialist who cared, to stumble on the idiosyncratic and the deep. A web reorganised into a summarisation layer over a shrinking base of primary material offers the reader something thinner: a smooth, confident, homogenised answer, increasingly assembled from the machine-shaped remnants of what the incentives have not yet killed.
You do not notice the loss as an event, because nothing is switched off on a particular Tuesday. You notice it, if at all, as a slow narrowing — the specialist blog that quietly stopped updating, the forum that went read-only, the guide that was never written because writing it no longer led anywhere. The answer box looks richer than ever while the ground it draws from gets poorer, and the reader who only ever sees the box has no way to perceive the depletion underneath. That is the quietly tragic shape of it: a convenience that feels like more while delivering, over time, less — and takes from the reader the very thing that made the web worth summarising in the first place.
What makes the situation genuinely awkward is that the platforms know all of this. The contradiction between summarising the web and starving it is not a subtle one, and the people running these systems are not naive about supply. The licensing deals, the publisher partnerships, the careful language about “supporting the ecosystem” — these are the sounds of an industry that understands the problem perfectly well and would prefer to be seen managing it than to actually resolve it, because resolving it means giving back the traffic that the answer box exists to capture. Awareness is not the missing ingredient. Willingness is, and willingness runs directly against the business model.
None of these are technically hard. They are commercially unwelcome, because the entire appeal of the answer box is that it keeps the user on Google's page and off everyone else's. That is the part worth saying plainly: the open web is not dying of natural causes. It is being quietly reorganised so that the value flows upward, and the reorganisation is being sold to users as convenience. It is convenient. So is spending down a savings account.
Finally someone said it. I cancelled my sub last week for exactly this reason.