Meta Appears to Be Building a Web Search Index for Its AI, Internal Sources and Crawler Docs Show

Meta Appears to Be Building a Web Search Index for Its AI, Internal Sources and Crawler Docs Show Meta Platforms appears to be constructing its own web search index, a move that would reduce the compa...

Meta Appears to Be Building a Web Search Index for Its AI, Internal Sources and Crawler Docs Show

Meta Platforms appears to be constructing its own web search index, a move that would reduce the company's dependence on Google for AI-powered search and put it in direct competition with existing search and answer engines. Developer Pieter Levels wrote on X that Meta employees had privately informed him of the project, and Meta's own published crawler documentation supports the claim.

The details

Levels, known for building solo software ventures like Nomad List and Photo AI, posted on August 7, 2026, that Meta staff had messaged him privately about the initiative. "Meta is ALLEGEDLY building their own Google search engine, so that if their AI does a web search it doesn't end up at Google, as Google could then use it for THEIR training," Levels wrote. "So they want their own web index that they will then use as their own Meta search engine for their AI." Levels noted he was sharing the information with the permission of his source.

In a follow-up post, Levels reported that Meta's crawlers had been hitting his websites aggressively. "Meta is doing heavy heavy heavy scraping on all my sites this week too (and seemingly everyone else's sites now)," he wrote. "So much so that I got load average alerts for it today on one VPS." The post accumulated 1.3 million views on X.

A respondent identifying himself as a former Meta employee offered additional context in the replies. "Worked at Meta for several years, and I'm quite familiar with the situation," wrote Vito Strokov. "They need more data to train a new class of AI models, the goal is to train 5-10T parameters model. Hence, they are trying to get whatever they can." The claim about a 5-10 trillion parameter model could not be independently verified.

Meta's own developer documentation supports the reports. The company maintains a public page titled "Meta Web Crawlers," last updated May 21, 2026, that lists five distinct crawler user agents. One of them, Meta-WebIndexer, is described explicitly as a crawler that "navigates the web to improve Meta AI search result quality for users." The documentation states that the crawler "analyzes online content to enhance the relevance and accuracy of Meta AI" and that allowing it in robots.txt "helps us cite and link to your content in Meta AI's responses."

Another crawler, Meta-ExternalAgent, "crawls the web for use cases such as training foundation AI models or improving products by indexing content directly." A third, Meta-ExternalFetcher, "fetches individual links at a user's request and supports product functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users." The documentation notes that Meta-ExternalFetcher may bypass robots.txt rules entirely.

The crawler descriptions, read together, describe infrastructure for building and maintaining a comprehensive web index. The Meta-WebIndexer documentation specifically frames the effort around powering Meta AI's search and citation capabilities, matching the insider account Levels shared.

Facebook has explored web search for over a decade. The company partnered with Bing in 2013 to power web search results within Facebook, only to end the arrangement in 2014. Since then, Meta has launched Meta AI across Facebook, Instagram, WhatsApp, and Messenger, and has open-sourced large language models in the Llama family, which have become widely used in the open-source AI community.

Why this matters for SEO and AI visibility

A Meta-built web index would introduce a new major player in the search ecosystem. Meta AI already reaches billions of users through its family of apps, and if the company can serve AI-generated answers with citations drawn from its own crawl rather than Google's or Bing's index, websites would need to optimize for yet another crawler and ranking system.

The Meta-WebIndexer crawler's documentation explicitly frames crawling as a prerequisite for being cited in Meta AI responses, which mirrors how Google's AI Overviews and Bing's Copilot cite web sources. Sites that block Meta's crawlers in robots.txt risk losing visibility in Meta AI's answers. SEO and GEO teams will need to decide whether to allow Meta-WebIndexer, Meta-ExternalAgent, and Meta-ExternalFetcher, each of which serves a different purpose and carries different implications for data usage.

The aggressive crawling Levels reported also raises infrastructure concerns. If Meta scales its web indexing operation to Google-level volumes, smaller publishers could face increased bandwidth and server costs. Meta's documentation recommends allowing crawlers via robots.txt and says the company honors standard robots directives, but Meta-ExternalFetcher may bypass those rules when fetching links at a user's request.

What to watch next

Meta has not publicly confirmed that it is building a standalone search engine. The company may frame the effort as infrastructure for Meta AI rather than a consumer-facing search product. Key signals to monitor include: expanded Meta-WebIndexer crawling volume, new Meta AI features that cite web sources, any announcement of a Meta search product, and how Google responds to a potential competitor building its own web index at scale. The robots.txt decisions publishers make in the coming months regarding Meta's crawlers will shape which content appears in Meta AI's answers.

Sources

  • Pieter Levels (@levelsio) on X, August 7, 2026: https://x.com/levelsio/status/2085467097405247945
  • Meta Web Crawlers documentation, Meta for Developers (updated May 21, 2026): https://developers.facebook.com/docs/sharing/webmasters/web-crawlers
  • Vito Strokov (@vitostrokov) reply on X, August 6, 2026: https://x.com/levelsio/status/2085467097405247945

Explore this topic

Keep following the same growth thread