ChatGPT Runs Its Own Search Index Family Called Labrador, Research Reveals
OpenAI's ChatGPT does not simply relay results from Google or Bing. According to new research published September 4, 2026, ChatGPT operates a multi-layered retrieval system internally called Labrador — a family of at least 12 specialized indexes covering general web, PDF, YouTube, news, arXiv, Wikipedia, local businesses, finance, legal, medical, shopping, and images. Most of these indexes were never publicly documented by OpenAI.
The finding fundamentally challenges the assumption that has dominated SEO and GEO (Generative Engine Optimization) strategy for the past two years: that ChatGPT is essentially a front end for existing search engines. The reality, according to the research, is considerably more complex — and has significant implications for how practitioners should think about content discoverability in AI search.
Background and context
When ChatGPT launched web browsing capabilities, the natural assumption was that it relied entirely on Microsoft's Bing or Google's search infrastructure. OpenAI's public communications have been deliberately vague, referring only to "third-party search providers" without naming them or describing the underlying architecture.
That ambiguity served a purpose. It allowed the SEO industry to apply existing Google optimization playbooks to ChatGPT, treating it as another search engine with a crawler, an index, and ranking algorithms. But the Labrador research reveals that ChatGPT's retrieval stack is not a single pipeline — it is a layered system with distinct components serving different functions, each with its own data sources, freshness characteristics, and blind spots.
The revelation also aligns with testimony from Nick Turley, Head of ChatGPT, who stated under oath in a 2025 US antitrust hearing that building an in-house index was OpenAI's plan from the beginning. Turley confirmed that OpenAI sought access to Google's data not as a permanent dependency but for quality benchmarking and to improve their own index. That testimony, previously treated as aspirational, now appears to describe an operational reality.
What exactly is Labrador
SEO analyst and researcher Tomek Rudzki, who has extensively documented AI search behavior, published findings showing that Labrador is not a monolithic index but a family of specialized indexes. Each serves a different content vertical:
- General web: The broadest index, covering standard web pages
- PDF: A dedicated index for document files
- YouTube: Video content and transcripts
- News: Time-sensitive journalism and reporting
- arXiv: Academic preprints and research papers
- Wikipedia: Encyclopedic reference content
- Local: Business listings and geographic data
- Finance: Market data, financial reports, and economic information
- Legal: Court documents, regulations, and legal analysis
- Medical: Health information and clinical research
- Shopping: Product listings and e-commerce data
- Images: Visual content and image metadata
Most of these specialized indexes were never mentioned in any OpenAI documentation. Their existence suggests a level of vertical specialization that mirrors — and in some cases exceeds — what traditional search engines offer through their own vertical search products.
The three-layer retrieval architecture
Independent research from RESONEO, a team that built a Chrome extension to capture ChatGPT's raw data stream, provides additional technical detail about how Labrador operates in practice. Their analysis of 1,200 ChatGPT answers, 88,000 search results, and 26,900 distinct pages revealed three distinct layers in the retrieval stack:
Layer 1: Discovery index. This is the primary retrieval layer — the system that finds candidate pages for a given query. It determines which pages enter the consideration set.
Layer 2: Reading cache. ChatGPT maintains a cache of full page copies it has previously fetched. This explains why some citations reference content that may no longer be live on the original URL, or why certain pages appear in responses even when they would be slow to fetch in real time.
Layer 3: Live fetches. A small subset of pages are opened live during the response generation process. This layer is constrained by latency — the researchers found that 58% of pages take more than 0.8 seconds to load, which is too slow for real-time inclusion in a conversational response.
Until July 21, 2026, ChatGPT's data stream included a field called result_source that identified which internal pipeline produced each search result. Four values appeared consistently: labrador, bright, oxylabs, and serp. None of these is documented by OpenAI. The field then vanished from the data stream overnight across all monitored accounts.
The RESONEO team now reidentifies each pipeline by its formatting signatures — snippet length, title shape, and other structural markers — using a classifier that achieves approximately 98% accuracy.
Where the data comes from
The research identifies multiple data sources feeding into Labrador:
- OpenAI's own crawler: The company operates its own web crawling infrastructure
- Bright Data: A web intelligence platform that appears to scrape Google search results pages directly
- Oxylabs: A proxy and scraping service that appears to target Google News verticals specifically
- Yelp and TripAdvisor: Local business data from established review platforms
- Web IQ: A web data provider (Microsoft has confirmed using this service)
This multi-source approach explains why ChatGPT's citations sometimes appear to come from Google results even when the underlying content was fetched through alternative channels. The system is not simply proxying Google's index — it is aggregating data from multiple providers, caching it, and then retrieving from cache when real-time fetching would be too slow.
What this means for SEO and GEO
The Labrador revelation has immediate practical implications for practitioners:
1. You are not optimizing for one system. If ChatGPT operates 12 specialized indexes, then content discoverability depends on which index serves your vertical. A medical website and a local business are subject to completely different retrieval rules, data sources, and freshness requirements.
2. Cache matters as much as crawl. If ChatGPT relies heavily on its reading cache, then being fetched once and cached is more valuable than being crawled repeatedly. This inverts the traditional SEO emphasis on crawl frequency. Content that is fetched, cached, and then updated on the original URL may not be reflected in ChatGPT's responses until the cache refreshes.
3. Latency is a ranking factor. The 0.8-second threshold for live fetching means that slow-loading pages are effectively invisible to the live retrieval layer. They can only appear in responses if they were previously cached. Page speed, already important for Google ranking, may be even more critical for AI search visibility.
4. Vertical specialization changes the game. If ChatGPT maintains separate indexes for legal, medical, financial, and other verticals, then authority signals and ranking factors may differ by vertical. A site that ranks well in the general web index may not rank well in the specialized legal or medical indexes.
5. Third-party data providers are part of the pipeline. The use of Bright Data, Oxylabs, and other scraping services means that ChatGPT's view of the web is mediated by intermediaries. How these services handle robots.txt, caching policies, and content freshness directly affects what ChatGPT sees.
Industry reaction
SEO expert Glenn Gabe, who has extensively tracked AI search behavior, offered a pointed assessment: the SEO industry has been optimizing for a "non-existent search engine." The assumption that ChatGPT simply uses Google or Bing has led practitioners to apply Google-centric optimization strategies to a system that operates on fundamentally different principles.
The revelation that ChatGPT maintains its own indexes — and that these indexes are specialized by vertical — suggests that GEO strategy needs to be rethought from first principles. Rather than asking "how do I rank in ChatGPT," practitioners may need to ask "which of ChatGPT's 12 indexes serves my vertical, and what are its specific ranking factors?"
What has not been confirmed
Several important questions remain unanswered:
- How are the specialized indexes ranked? Labrador may maintain separate indexes, but it is not clear whether they use different ranking algorithms or the same algorithm with vertical-specific training data.
- What is the cache refresh cycle? The research confirms that caching is important, but the frequency of cache updates — whether hours, days, or weeks — has not been disclosed.
- How does OpenAI handle conflicting data? When Bright Data's scrape of Google disagrees with OpenAI's own crawler, which source takes precedence?
- Are all 12 indexes operational for all users? The research identifies the indexes by name, but it is not clear whether all are active in production or whether some are experimental.
- What role does user feedback play? ChatGPT allows users to regenerate responses or provide feedback. It is not clear whether this feedback influences Labrador's ranking algorithms.
What to watch next
OpenAI's job postings provide one signal. Recent listings mention "exabyte-scale database systems," suggesting that the company is investing heavily in storage and retrieval infrastructure at a scale that goes well beyond a simple proxy for Google or Bing.
The disappearance of the result_source field from ChatGPT's data stream on July 21 suggests that OpenAI is aware that researchers are reverse-engineering its retrieval architecture and is taking steps to obscure the internal workings. This opacity makes independent verification more difficult.
For practitioners, the immediate next step is to monitor which types of content appear in ChatGPT responses for their vertical and to test whether changes to their content affect AI search visibility. The Labrador research suggests that the relationship between content and AI search citations is more complex than a simple ranking algorithm — it involves caching, latency, vertical specialization, and third-party data providers.
Historical precedent
This is not the first time a major technology company has obscured its search infrastructure. Google's early search algorithm was famously opaque, and the SEO industry spent decades reverse-engineering ranking factors through observation and experimentation. The difference with ChatGPT is the speed of change — model updates, cache refreshes, and index changes can happen without the gradual evolution seen in traditional search engines.
The 2025 antitrust testimony from Nick Turley now appears prescient. OpenAI always planned to build its own index. The question was never whether they would, but when they would reveal that they had.
Sources
- Rudzki, Tomek. "ChatGPT runs its own index." LinkedIn post, September 4, 2026. https://www.linkedin.com/posts/tomekrudzki_chatgpt-runs-its-own-index-chatgpts-activity-7501622620673667072-TRQ-
- de Segonzac, Olivier (RESONEO). "Inside ChatGPT's retrieval stack: The index, cache, and pages it actually reads." Research based on analysis of 1,200 ChatGPT answers and 88,000 search results, published August 2026.
- Turley, Nick (Head of ChatGPT, OpenAI). Testimony before US Congress, 2025 antitrust hearing. Public government record.




