Perplexity Cites Its Sources. Two Audits Say the Sources Don’t Check Out.
The answer engine that sells itself on showing its sources has a sourcing problem — and two research teams measured it in the same week.
Perplexity’s pitch has always been the footnote. Ask it a question and you don’t just get an answer the way a chatbot gives you one — you get an answer with little numbered citations hanging off it, each one promising that a real page somewhere backs up what you just read. That is a genuinely good idea, and it is the entire reason a lot of people trust an answer engine over a raw model: you are told, explicitly, that you can check.
This week two independent research teams decided to actually check. Both published on the same day, 2 September 2026. One counted how often Perplexity’s citations point at a page that supports the claim attached to them. The other traced where the citations come from in the first place. Neither set out to prove the answers were wrong. Both came back with the same uncomfortable finding from different directions: the citation — the one feature meant to let you verify an answer without trusting the machine — is frequently the part you cannot check.
This matters because “sourced” does an enormous amount of persuasive work. An answer with footnotes feels like journalism; an answer without them feels like a guess. If the footnotes don’t hold up, the confidence they buy is unearned — and that is a subtler problem than a model simply making something up.
The footnote that won’t open
The first audit, from Haus Research, took a deliberately boring approach: 310 factual questions about 210 technology companies — who runs them, how many people they employ, when they were founded — and then examined every citation Perplexity attached to a sentence stating a figure. That came to 1,826 citations. The test was not “is the number right,” but the narrower, more basic one: does the page you pointed me to actually contain this number?
It often did not. Haus found that 34.7% of the citations pointed to pages that either would not open or contained none of the figures they were cited for. Scored a stricter way — per claim, counting a claim as failed only when none of its cited pages supported it — 14.4% still failed outright. The weakest category was the most human one: questions about who runs a company passed just 44.3% of the time. Headcount held up best, at 82.0%. And a quarter of the cited URLs — 25.1% — had never once been captured by the Wayback Machine, which means that if the page changes or vanishes, there is no independent record that it ever said what it was cited as saying.
It helps to picture the mechanism, because none of this requires bad faith. The model retrieves some pages, writes a fluent sentence containing a figure, and staples a citation marker to it. Nothing in that pipeline guarantees that the figure in the sentence appears on the page behind the marker — the citation is generated alongside the claim, not derived from a line someone actually read. Unless a human clicks, the marker is decoration.
Haus was careful about what this does and doesn’t show. “We are not measuring whether Perplexity is right,” the report notes; “we are measuring whether the thing it offers as proof functions as proof.” That is the right frame. A citation is a promise that someone can retrace your steps. When a third of them lead to a locked door, the promise is broken whether or not the answer behind it happened to be correct.
The sources that were built to be cited
The second audit, from Trellner Research (report TR-2026-009), asked a different question: not whether the citations open, but where they come from. It ran 380 buyer-intent software categories — the “best CRM,” “best password manager,” “best project-management tool” queries that convert into sales — through Perplexity’s Sonar and Sonar Pro models, 760 API calls in all, asking each time for a ranked top five with homepage domains.
The citations that came back were drawn from 2,055 different domains, and a startling share of them were obscure: 59.8% pointed to domains ranked worse than 100,000th on the Tranco list of popular sites, and 23.4% to domains outside the top million entirely. Buried in that long tail was a pattern. Three linked domains — wifitalents.com, worldmetrics.org and gitnux.org, which Trellner notes “all three delegate DNS to the same pair of Cloudflare nameservers” — had between them generated 215,128 machine-made “best software” pages. That single network became the third-largest source of evidence behind Perplexity’s recommendations.
The detail that should worry anyone selling or buying software is who these pages are written for. “These pages are addressed, in their titles and descriptions, to the software that reads them,” Trellner writes — not to human readers, but to the retrieval systems of AI answer engines. This is the new search-engine optimisation: not gaming Google’s ten blue links, but manufacturing citeable-looking authority at industrial scale for models that grab a handful of sources and summarise them. We have watched this pattern build for a while — it is the same dynamic driving the death of the open web through AI overviews and why AI search is making Google worse. Trellner’s dry summary is the line to remember: “a vendor’s own listicles about markets it does not operate in became the third-largest evidence base.”
Why “sourced” feels safer than it is
Put the two findings together and you get the shape of the problem. One team found that many citations don’t support their claim; the other found that many of the sources being cited are junk built for exactly this purpose. In both cases the answer arrives wearing the costume of verification — numbered, footnoted, apparently checkable — while the substance underneath is thin.
This is worse, in one specific way, than a plain hallucination. When a model invents a fact with no source, a sceptical reader has a chance of noticing. When it invents — or borrows — the appearance of a source, it disarms exactly the instinct that would have caught the error. The citation says “don’t take my word for it,” and most people, reasonably, don’t click. It is a close cousin of the confidence problem we keep coming back to in AI hallucinations are still not solved: the machine is never more convincing than when it is quietly wrong.
It is also not unique to Perplexity, even if Perplexity is the one being measured here. The habit of an answer engine gesturing at sources it hasn’t really consulted is something we documented when Gemini often won’t search the web — and won’t tell you it didn’t. The reason these audits single out Perplexity is precisely that Perplexity made the citation its brand. If you sell the footnote as the product, the footnote is fair game.
To be fair: what the audits don’t prove
Two honest caveats, because the gap between promise and product cuts both ways. First, both audits ran against Perplexity’s Sonar and Sonar Pro models through the API, not the polished consumer app most people use, with its live-rendered source cards. It is entirely possible the flagship product does better, and Perplexity would be within its rights to say so.
But that defence only goes so far. The API models are what a large number of downstream “ask-our-docs” and research features are built on, so their behaviour is not a niche case. They draw on the same retrieval-and-ranking machinery as the main product. And the manufactured-page problem is upstream of any interface: if the index that feeds the model rewards content farms addressed to machines, no amount of tidy front-end design fixes what is being retrieved. A nicer card around a bad source is still a bad source.
Second, none of this makes AI search useless, and it would be its own kind of hype to pretend otherwise. Citing sources at all is better than the confident, sourceless paragraph a raw chatbot hands you; the audits are only possible because Perplexity shows its working. The criticism is not that the idea is bad. It is that the promise has to be kept, and right now, by two independent measurements, it often isn’t.
What to do when the answer comes with footnotes
Until the sourcing catches up with the marketing, the safest posture is to treat a citation as a lead, not a verdict. Concretely:
- Click the number. If the linked page won’t open, or you can’t find the exact figure on it, discount the claim — roughly a third of the time, on Haus’s count, that is what you’ll find.
- Be most sceptical of superlatives. “Best X software” answers are exactly where the manufactured listicle pages cluster, because that is where the buying intent — and the affiliate money — lives.
- Prefer primary pages over aggregators. A citation to a company’s own documentation, a regulator, or a named publication is worth more than one to a domain you’ve never heard of ranked below the top million sites.
- Verify anything that costs you. A price, a headcount, who runs a company, a security claim — confirm it against a source you would have trusted before AI search existed.
The footnote was a real advance. It promised to turn an AI answer from something you believe into something you can check. The disappointing news from this week is that, for now, the promise and the product have come apart: the citation is offered as proof, but a large share of the time it doesn’t function as proof. That is not a reason to stop using AI search. It is a reason to click the number — and to trust the answer a little less until the source underneath it earns it.
Frequently asked questions
Did a third of Perplexity’s answers turn out to be wrong?
No — that is not what the audits measured. Haus Research found that 34.7% of the citations Perplexity attached to factual claims either wouldn’t open or didn’t contain the figure they were cited for; it did not test whether the underlying answer was correct. The distinction matters: an answer can be right while its cited proof is unverifiable, and the whole point of a citation is to let you check without taking the machine’s word for it.
What were the sites that Perplexity kept citing?
Trellner Research named three linked domains — wifitalents.com, worldmetrics.org and gitnux.org — that together had published 215,128 machine-generated ‘best software’ listicle pages. The three delegate DNS to the same pair of Cloudflare nameservers, suggesting common control, and their pages are written to be read by AI models rather than people. In Trellner’s test of 380 software categories, this network became the third-largest source of citations behind Perplexity’s ranked recommendations.
Does this affect the Perplexity app I use, or just the API?
Both audits ran against Perplexity’s Sonar and Sonar Pro models through the API rather than the consumer Pro app, which is a fair limitation to note. But the API models power a large number of third-party ‘answer’ features, and they draw on the same retrieval-and-citation machinery as the main product. The safest reading is that the sourcing weakness is a property of how these systems select and attach citations, not a quirk of one interface.
How is this different from a hallucination?
A hallucination is the model inventing a fact. This is subtler and arguably harder to catch: the fact may be real, but the citation offered as proof doesn’t support it, won’t load, or traces back to content generated specifically to be cited. Because the answer looks sourced, it invites more trust than an unsourced one, even when the source is the weakest link.
What can I actually do about it?
Treat the citation as a lead, not a verdict: click through, and if the page won’t open or doesn’t contain the exact figure, discount the claim. Be most sceptical of superlative ‘best X software’ answers, which is exactly where manufactured listicle pages cluster. And for anything that matters — a price, a headcount, who runs a company — confirm it against a primary source you would have trusted before AI search existed.
Sources
- Haus Research — “Perplexity Citation Audit”: 34.7% of citations pointed to pages that wouldn’t open or contained none of the figures cited (2 Sep 2026) — Haus Research
- Trellner Research — Report TR-2026-009, “Manufactured Sources Behind AI Recommendations”: three domains generated 215,128 ‘best software’ pages that became the third-largest evidence base (2 Sep 2026) — Trellner Research
