ChatGPT's Fetch Bot Reaches Pages That Disallow It — OpenAI Says Robots.txt May Not Apply

TollBit's H1 2026 report shows ChatGPT-User is the most-blocked AI page-fetcher yet reaches disallowed pages on roughly half of EU sites that listed it. OpenAI documentation says robots.txt rules may not apply because a person initiated the request.

ChatGPT's page-fetching bot is blocked by more sites than any other AI bot of its kind, and it reaches pages those sites explicitly disallowed on roughly half of European domains that listed it. OpenAI's own documentation states that robots.txt rules may not apply to the bot because a person initiated the request.

What the data shows

TollBit's State of the Bots report for the first half of 2026 tracked how AI page-fetching agents interact with publisher robots.txt files. On European sites monitored by the report, about 15% of identified AI page-fetchers reached URLs that site operators had marked as disallowed.

The bypasses concentrate on a small number of agents. ChatGPT-User, Bytespider, and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them in their robots.txt. Among these, ChatGPT-User reached the most sites.

Newer page-fetching agents face much lower block rates. Only 9% of European websites disallow Claude-User, compared with 26% in North America. Perplexity-User sits at 13% in Europe versus 26% in North America. Most newer agents have disallow rates in the single digits across Europe.

What OpenAI says about the rule

OpenAI's crawler documentation describes ChatGPT-User as the agent that visits a page when a ChatGPT user asks a question. Because those actions are user-initiated, the documentation states that robots.txt rules may not apply.

Perplexity takes the same position for Perplexity-User, saying it generally ignores robots.txt for the same reason. Anthropic has a different view and states that all three of its bots respect the file.

TollBit treats any request to a disallowed URL as a bypass regardless of what the operator claims about the legal basis for the request.

What this means for site operators

According to OpenAI's documentation, the agent responsible for deciding whether a site appears in ChatGPT search results is OAI-SearchBot, not ChatGPT-User. Sites that block both agents to prevent AI traffic have blocked the visibility decision-maker and kept a fetching control that carries a documented carve-out.

Server logs or CDN records show what actually reached a site. A robots.txt file shows only what the operator asked for.

Cloudflare's network-layer enforcement

Cloudflare is moving AI bot management to the network layer. Starting September 15, new domains added to Cloudflare will have Training and Agent crawlers blocked by default on pages with ads, while Search crawlers remain allowed. When Cloudflare recognizes a bot, compliance is no longer left to the crawler itself.

Whether the user-initiated argument survives network-layer enforcement is the open question. The argument rests on the distinction between requesting a page and a crawler taking it. All major AI assistants now fetch pages this way.

What has not been confirmed

The TollBit report covers European and North American sites monitored by the platform. It does not establish global bypass rates. OpenAI has not publicly responded to the report's findings beyond its existing documentation. The September 15 Cloudflare change applies to new domains; existing domains may follow a different timeline.

Sources

  • TollBit, "State of the Bots: 2026 Q1 & Q2, The Bad Bots" — https://tollbit.com/state-of-the-bots/
  • Search Engine Journal, "OpenAI Says Robots.txt May Not Apply To ChatGPT's Fetch Bot" (Matt G. Southern, August 2026) — https://www.searchenginejournal.com/openai-says-robots-txt-may-not-apply-to-chatgpts-fetch-bot/585864/

Explore this topic

Keep following the same growth thread