215,128 Manufactured 'Best Software' Pages Got Cited by Perplexity: What It Means for GEO

Key takeaways

Three domains generated 215,128 machine-made 'best software' pages and earned 7,534 Perplexity citations. Here's what the report actually proved, what it didn't, and how legit sites should respond.

Three domains that did not exist before December 2023 published 215,128 machine-generated "best software" buying guides. In a September 2, 2026 measurement, Perplexity cited those pages hundreds of times, and 59.8% of all the citations the study collected pointed at domains ranked below the 100,000th most popular site on the web. The short version for GEO teams: AI citation manipulation is real, it is detectable, and the response is not a new tactic — it is tighter evidence hygiene and measurement.

The study, precisely scoped

Trellner Research's report TR-2026-009 (September 2, 2026) is descriptive, not prescriptive. It set out to answer one question: which documents make up the evidence base Perplexity cites? The method: 760 calls to Perplexity's sonar and sonar-pro models over 380 software categories, collecting 7,534 citations across 2,055 distinct domains. The dataset, scripts, and method are released under CC BY 4.0.

Scope matters, because the report is easy to over-read. Only Perplexity was measured, through OpenRouter. Google was explicitly excluded, and the report makes no claims about ChatGPT, Gemini, Copilot, or Google's AI Mode. Anyone summarizing this as "AI search is overrun by fake pages" is extending the evidence beyond what it shows — including parts of the industry commentary that circulated around the release.

Three domains, 215,128 buying guides

The three domains are wifitalents.com, worldmetrics.org, and gitnux.org:

Domain

Generated /best/<category>-software/ pages

wifitalents.com

70,731

worldmetrics.org

71,684

gitnux.org

72,713

Total

215,128

As the report notes dryly, "there are not 215,128 software categories." None of the three domains existed before December 2023. The pages follow one template, and several carry unrendered template variables — two pages still read "Within the next 26 days," placeholder text that never got filled in.

The report stops short of naming an operator, but the evidence for one operation is laid out: shared NameCheap registrations, identical Cloudflare nameservers across all three brands plus a fourth domain (zipdo.co), identical page templates, and six cross-promotional blog posts in which each brand endorses the sibling brands. Worldmetrics also monetizes the traffic it captures, selling custom research "from €5,000" and vendor selection work "from €2,500."

The pages were written for retrieval software, not readers

The most striking detail is that two of the three domains announce what they are. Both worldmetrics.org and gitnux.org return an HTML title of "Facts & Grounding Page" and a meta description that positions the page as a machine-readable record of verified facts about the brand, explicitly addressed to retrieval software rather than to humans.

That framing matters beyond this one case. The pages are not trying to fool a human judge into thinking they are editorial content. They are built to look like clean, fact-dense documents to an automated retrieval pipeline. Perplexity's own documentation has encouraged brands to make information easy to extract — "grounding" is a legitimate GEO concept. This operation took that concept and industrialized it without the underlying substance: invented staff bylines (each page credits three distinct people, nine different names for a single question across three brands), template artifacts, and no verifiable authors anywhere.

Flow diagram of how a manufactured grounding page got cited: mass listicle template, self-declared Facts and Grounding Page metadata, retrieved by the crawler, then cited as a source in an AI answer

What the evidence base actually looked like

The citation-quality numbers are the core of the report:

Measure

Result

Citations pointing at domains ranked worse than #100,000 (Tranco top 1M)

59.8%

Citations pointing at domains outside the top 1M entirely

23.4%

Median cited-domain rank

71,611

Categories where sonar and sonar-pro returned byte-identical citation lists

289 of 380

Bar chart of Perplexity citation quality measures: 59.8 percent of cited domains ranked below position 100,000, 23.4 percent outside the top 1 million, median cited-domain rank 71,611, and identical citation lists across model tiers in 289 of 380 categories

The near-identical results across the two model tiers (Jaccard overlap 0.898) suggest the tiers share essentially one retrieval layer — which means measuring one tier tells you most of what you need to know about the other.

The models did not just cite weak domains; they recommended broken ones. Of 1,502 vendor homepages supplied by the models, 17 (1.1%) were dead or unreachable — including graphiql.com offered for GraphiQL, todo.com offered for Microsoft To Do, and aquasecurity.io offered for Trivy. Worse, in two categories the recommended domain was an outright redirect trap: for research data platforms, sonar-pro correctly returned datadryad.org while sonar returned dryad.co — which redirected to a gambling portal called BIGSLOT288; for data-quality tools, sonar returned montecarlodata.com while sonar-pro returned montecarlo.com, which now belongs to a Monaco hotel-and-casino operator.

There is also a quality-control finding buried in the same data: the three manufactured brands disagree with each other. For "project estimation software," Worldmetrics recommended Float, Scoro, Teamwork.com, Procore, and Wrike; WifiTalents recommended Float, Scoro, Teamwork.com, Buildertrend, and Apropo; Gitnux recommended Saviom, Mosaic, Buildertrend, Float, and Teamwork.com. Gitnux's winner appears nowhere in Worldmetrics' top five. Even a successful manipulation produces inconsistent answers — a reminder that the "recommendation" is only as coherent as the retrieval stack behind it.

Legitimate content floats in the same pool

The report's most-cited domains list is a useful sanity check: g2.com (291 citations), reddit.com (261), gartner.com (158), zapier.com (82), capterra.com (68), linkedin.com (67). Wikipedia, by contrast, was cited just 3 times across all 380 categories.

Two lessons sit side by side here. First, earned authority still concentrates in the expected places — the review and community hubs that AI engines treat as defaults. Second, ordinary content marketing can still win: guideflow.com, a demo vendor's marketing blog, became the third most-cited domain (194 citations across 96 categories, ahead of Gartner) with no manufactured network behind it — just consistent, category-specific content that answered the queries directly. The report includes it as a control case showing what a single real brand can do against 215,000 generated pages.

What the study does — and doesn't — prove

Four limits matter for anyone citing this report later:

  • Perplexity only. No claim here extends to ChatGPT, Gemini, Copilot, or Google AI Mode.
  • One-day snapshot, one prompt per category. No repeat sampling, no prompt variation; the 380 categories are proprietary and weighted toward niche verticals, which surfaces more long-tail sources than a broad consumer query set would.
  • Automated fetching. Pages were fetched through datacentre proxies with a research user-agent; four of the seventeen "unreachable" domains (including nasdaq.com and solidworks.com) are plainly alive and simply did not answer automated requests.
  • Inference, not proof, on ownership. The common operation is inferred from registrations, nameservers, and templates. No operator is named. And the report did not test whether removing these sources would change actual recommendations.

Treat the 59.8% and 23.4% figures as lower bounds on the noise in one engine's long-tail retrieval — not as a measurement of AI search overall.

What it means for GEO teams

Four implications follow, and none of them is "go generate more pages."

A citation is not an endorsement, so stop counting it as one. If a retrieval pipeline will happily cite a domain ranked 700,000th that did not exist two years ago, then a citation in an AI answer is weak evidence of authority by itself. Citation reports that count only appearances will mislead you. Add citation quality checks — who else is cited alongside you, what those domains rank, whether the answer's facts actually match your page — or you will celebrate the same signal the manufactured pages are buying.

You can be cited wrongly, and you will not see it in rank tracking. A redirect trap that sent Perplexity to a casino portal is the extreme case; the ordinary case is an AI answer that paraphrases your content behind a competitor's name, or cites your stats without your brand. Monitoring what AI surfaces say about your entity is separate from monitoring where you rank.

Earned corroboration still concentrates in a handful of verified hubs. G2, Reddit, Gartner, Capterra, and LinkedIn absorbed most of the legitimate citations in the study. The brands that win AI answers at scale are the ones whose identity and claims are corroborated across those hubs — which is an entity problem, not a content-volume problem.

If your program looks like this operation, even a little, assume detection is trivial. The pages self-identified in their own metadata ("Facts & Grounding Page"), carried unrendered template variables, and listed staff who did not exist. If your own "best X" program relies on mass template output with thin authorship, this report is the template for how your program will be exposed. The durable alternative — building a page a human expert would stand behind, with named, verifiable authorship and genuinely divergent recommendations — is also the one that survives the next retrieval-model update. It is the difference between your page being the source of an answer and your page being the artifact the answer's flaws get traced back to.

FAQ

Is this happening in ChatGPT, Gemini, or Google AI Mode too?

Unknown. The study deliberately measured Perplexity only, and its authors make no claims about other engines. The mechanisms it documents (mass template output, self-describing grounding pages, retrieval-oriented metadata) are engine-agnostic in principle, but "in principle" is not evidence.

Was this a rogue SEO agency?

The report does not name an operator. It documents shared registrations, nameservers, templates, and cross-promotion across four domains, and notes that worldmetrics.org sells paid research and vendor-selection services — a plausible monetization model — but ownership remains inferred.

Does this mean Perplexity is unreliable for research?

It means long-tail categories deserve a second pass. The same study shows the engines' default behavior is still dominated by legitimate hubs (G2, Reddit, Gartner). The failure mode is concentrated in niche categories where the retrieval layer reaches for whatever structured-looking document exists — which is where a quick cross-check against the vendor's own site pays off.

Should I stop building comparison and "best X" content?

No — but build it like the report's control case, guideflow.com, not like the manufactured brands: real authorship, real testing, named sources, and recommendations that survive a human expert's scrutiny. Comparison content is a core AI-search surface; the differentiator is whether a reader (or a retrieval system) can verify who made the call and why. If your brand already publishes best-of lists, our earlier breakdown of why a best-X article can make AI answers recommend your competitors is the natural companion to this one.

Author: Adrian Cole, Analyst of 1,000+ AI Search Results at Auspia. Adrian writes about how brands appear in ChatGPT, Perplexity, Gemini, and other answer surfaces, and what the evidence says about who gets cited and why.

Explore this topic

Keep following the same growth thread