Three Sites Generated 215,128 Pages to Feed Perplexity AI Recommendations, Trellner Research Finds

Key takeaways

Trellner Research documents the first large-scale case of content engineered for AI retrieval: three sites published 215,128 machine-generated pages that Perplexity AI cited in its grounded recommendations.

Three Sites Generated 215,128 Pages to Feed Perplexity AI Recommendations, Trellner Research Finds

A new report from Trellner Research has documented the first large-scale, empirically measured case of content specifically engineered to be retrieved by an AI search engine — and cited as evidence in its answers. Three websites under apparently common control published 215,128 machine-generated "best software" pages between them, and Perplexity AI's grounded recommendation engine cited them alongside established sources like Gartner and Reddit.

The findings, published September 2, 2026, reveal what researchers call "manufactured sources" — content built not for human readers but for the retrieval step that precedes an AI model's answer. The report provides the most detailed public evidence to date that generative engine optimization (GEO) spam is already an operational reality, not a theoretical future risk.

Background and context

The report arrives at a moment when B2B software buyers are fundamentally changing how they begin purchase research. According to a G2 survey of 1,076 B2B software buyers and decision-makers conducted in March 2026, 51% now begin their software research with an AI chatbot more often than with Google — up from 29% in April 2025. The same survey found that 71% of buyers rely on AI chatbots for software research, compared with 60% just seven months earlier.

The commercially significant finding is not the adoption rate itself. G2 reports that 69% of buyers chose a different software vendor than initially planned based on AI chatbot guidance, and that one in three (33%) purchased from a vendor they were not previously familiar with. An AI assistant is not merely reordering a consideration set the buyer already held — in a third of cases, it is introducing names that would not otherwise have entered the evaluation.

A separate Gartner survey of 645 B2B buyers between August and September 2025 found that 45% used generative AI, primarily to gather information on vendors and products, within an average of seven information sources used during a recent purchase. The convergence of these independent measurements confirms that the opening move in a significant and growing share of B2B purchase decisions now happens on a surface that produces no impression data, no query record, and no ranking position that vendors can inspect or optimize through traditional means.

What the research measured

Trellner Research tested Perplexity's grounded AI models — perplexity/sonar and perplexity/sonar-pro — across 380 buyer-intent categories ranging from "CRM software" to "museum collection management software." The researchers issued 760 API calls through OpenRouter, one prompt per category per model, each asking for a ranked top five with each product's official homepage domain. All 760 calls returned parseable answers with the URLs the models retrieved during grounding.

The run produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. Trellner then looked up every cited domain in the Tranco daily list for September 1, 2026, and in the Wayback Machine, and fetched every one of the 1,502 vendor homepages the models supplied to verify whether they still existed.

The headline finding: 59.8% of the 7,534 citations point at domains ranked worse than #100,000 in the Tranco top-1M list, and 23.4% point at domains that do not appear in the top million at all. The median Tranco rank of the 5,768 citations that do point at a ranked domain is 71,611.

The three "facts and grounding" sites

Three domains in the results — wifitalents.com (71 citations across 27 categories), worldmetrics.org (60 citations, 22 categories), and gitnux.org (50 citations, 23 categories) — account for 181 citations, 2.4% of the total, and appear together in 41 of the 380 categories tested.

Trellner's investigation found strong circumstantial evidence that all three operate under common control:

  • All three were registered through NameCheap between December 2023 and May 2024
  • All three delegate DNS to the same pair of Cloudflare nameservers (pam.ns.cloudflare.com and sean.ns.cloudflare.com)
  • All three run the same page template with identical navigation: Services, Market Data, Software Advice, Editorial Process, Company
  • Each maintains a blog of exactly six posts, and all eighteen posts are about the other brands in the set
  • A fourth brand, zipdo.co, sits on the same nameserver pair and gives its homepage the identical "Facts & Grounding Page" title

Their scale is the operational detail that separates them from ordinary content marketing. Their sitemaps list 103,578, 107,083, and 105,541 URLs respectively, of which 70,731, 71,684, and 72,713 are /best/<something>-software/ pages: 215,128 generated buying guides across three brands. There are not 215,128 software categories — the pages cover overlapping and invented categories at industrial scale.

The self-description is what makes them structurally unusual. Fetched on September 2, 2026, worldmetrics.org and gitnux.org both return an HTML title of the form "<Brand> — Facts & Grounding Page," with a meta description reading: "Verified facts about [Brand]: an independent market research company publishing industry statistics, custom research, and software Best Lists. Company, legal, methodology, and compliance details in one machine-readable record."

Grounding is not a term buyers use. It is the name of the step in which a retrieval system fetches documents to condition an answer on. A machine-readable record of verified facts about oneself is not a service to a human reader. These pages are addressed, in their titles and descriptions, to the software that reads them.

One template, three contradictory verdicts

Trellner fetched the same category page — "project estimation software" — from all three brands. Each page states its ranking in JSON-LD structured data, so it can be read without interpretation. The three sites produced three different rankings for the same question:

  • worldmetrics.org ranked Float first, followed by Scoro, Teamwork.com, Procore, and Wrike
  • wifitalents.com ranked Float first, followed by Scoro, Teamwork.com, Buildertrend, and Apropo
  • gitnux.org ranked Saviom first, followed by Mosaic, Buildertrend, Float, and Teamwork.com

Gitnux's winner (Saviom) does not appear in worldmetrics' top five at all. Each page carries three named staff authors — nine distinct people across the three sites for one question. Each page announces an editorial process; Gitnux labels its result "AI-verified · Expert reviewed." All three carry an unrendered template variable in the byline line, reading "Within the next 26 days" on two of them and "Within the next 40 days" on the third — a visible artifact of incomplete automation.

The vendor marketing blog that outranks Gartner

The report's third-largest cited source is not one of the three manufactured brands. guideflow.com, which sells interactive product demos and competes in none of the 380 categories tested, was cited 194 times across 96 of the 380 categories — a quarter of them — placing it third overall and ahead of Gartner (158 citations). Each citation pointed to a different URL: 96 distinct guideflow.com blog URLs, one per category.

Guideflow publishes a large content-marketing blog, as thousands of companies do. Its sitemap lists 3,351 blog URLs, 2,176 of them distinct posts. The measurement is not about deception — Guideflow does not pretend to be a review site. The finding is about what the retrieval layer does with vendor content: a vendor's own listicles about markets it does not operate in became the third-largest evidence base for a question about which product to buy.

What this means for GEO practitioners

The report has immediate implications for anyone working in generative engine optimization:

The retrieval layer is the new ranking layer. Traditional SEO optimizes for position in a sorted list. GEO must optimize for inclusion in the set of documents a model retrieves before generating an answer. The Trellner data shows that retrieval does not discriminate between established authorities and purpose-built content farms — it returns what matches the query pattern, regardless of provenance.

Scale beats authority in the retrieval step. The three manufactured sites, none of which existed before December 2023, collectively outrank Wikipedia (cited three times in 7,534 results) and compete directly with Gartner, Capterra, and Zapier. Domain authority, as traditionally measured, correlates poorly with retrieval success in this context.

Structured data is read without interpretation. All three manufactured sites publish their rankings in JSON-LD, which the models can extract directly. The content does not need to persuade a human editor or pass a quality review — it needs to be machine-parseable and topically aligned with the query.

The verification step is where quality still matters. Trellner's companion report notes that 94% of B2B buyers who used AI during a purchase said they fact-check its responses at least some of the time, and 69% prefer to validate AI-generated insights with sales representatives. The AI assistant's description is the first impression; the material a buyer finds when verifying it is what determines whether that impression survives. Vendors with thin, absent, or inconsistent public records lose at the verification step regardless of what the assistant initially said.

What has not been confirmed

The report is careful about its own limitations. The two Perplexity models tested (sonar and sonar-pro) are not independent measurements — they returned byte-identical citation lists in 289 of 380 categories, with a Jaccard overlap of 0.898. They share a retrieval layer and should be read as one search stack sampled twice.

The results cover Perplexity only. Trellner has not measured ChatGPT, Gemini, Copilot, or Google's AI Mode, and there is no reason to assume their retrieval mixes match. The 380 categories are the researchers' own construction, not a sample of what buyers actually ask, and a list weighted towards niche verticals will surface more long-tail sources than a list of common queries would.

Trellner has not shown that any of this changes the answers. They did not test whether removing these sources would produce different recommendations. The three manufactured sites and Guideflow may well name reasonable products. What was measured is which documents the evidence base is made of, not whether the final recommendations are worse for it.

Common control of the three brands is inferred from shared infrastructure and an identical template. Trellner does not know who operates them; none of the three names an owner.

What to watch next

The report raises questions that the industry will need to answer:

  • Will Perplexity or other grounded AI engines implement source quality filters that distinguish purpose-built content from organic editorial work? The current retrieval layer clearly does not.
  • Will the "facts and grounding" pattern spread beyond software into other high-value B2B categories where AI-assisted purchasing decisions are growing?
  • How will established review sites and directories respond when their citation share is being displaced by domains that did not exist two years ago?
  • Will search engines and AI providers develop provenance signals — similar to how Google's E-E-A-T framework attempts to assess creator expertise — that can surface in retrieval rankings?

The full dataset, including every citation, every recommendation, the Tranco and Wayback lookups, the vendor liveness checks, and the scripts that produced every figure, is published under CC BY 4.0 at Trellner's data repository.

Sources

Explore this topic

Keep following the same growth thread