ChatGPT’s page-retrieval bot encounters restrictions from more websites than any comparable AI bot currently in existence.
Notably, it has accessed forbidden pages on a greater number of sites than its counterparts. OpenAI contends that the standard protocols delineated in robots.txt may not be applicable, asserting that a request from an individual necessitates access to the page.
TollBit’s most recent report, the State of the Bots, provides statistical insights for the first half of 2026. This document further elucidates the behavioral patterns of various crawlers and their implications for your website.
Bypassing Restrictions
According to the findings from the report, around 15% of AI page-retrievers successfully navigated to URLs that European websites had expressly marked as disallowed.
This phenomenon primarily involves a select few agents. For instance, ChatGPT-User, Bytespider, and Youbot managed to access restricted pages on nearly half of the European sites that had explicitly flagged them. Among these agents, ChatGPT-User gained access to the most sites.
Disallowed Access
A significant number of recent page-fetching agents encounter minimal blocking. Only 9% of European websites impose restrictions on Claude-User, in stark contrast to 26% in North America. Similarly, Perplexity-User faces a disallowance rate of 13% compared to 26% for North American sites.
While most emerging agents have disallowance rates lingering in single digits across Europe, ChatGPT-User remains a notable exception to this trend.
OpenAI’s Stance on Restrictions
OpenAI’s documentation asserts that ChatGPT-User visits a page only upon user request, suggesting that due to the user-initiated nature of these actions, robots.txt rules may not be applicable.
Perplexity claims that Perplexity-User generally disregards the restrictions for similar reasons, while Anthropic adopts a differing approach, maintaining that all three of its bots adhere to these guidelines, as noted in our coverage from February.
TollBit considers any approach to a blocked URL as a circumvention, irrespective of the operator’s assurances.
Significance of These Findings
A disallow directive for ChatGPT-User is viewed as a request that, per OpenAI’s documentation, might not be enforced.
It is imperative to contemplate an additional facet of this issue. According to OpenAI’s guidelines, the entity responsible for determining a site’s visibility in ChatGPT search outcomes is designated as OAI-SearchBot, not ChatGPT-User.
Sites that prohibit both agents to mitigate AI traffic have relinquished one aspect of visibility while retaining a fetching control that includes an exception.
Server logs and CDN records accurately reflect the data that has been requested, whereas the file solely indicates what was solicited.
Future Developments
Cloudflare is making alterations to its management of crawler controls, transitioning decision-making to the network layer. For the bots it recognizes, compliance will no longer be contingent upon the crawler itself.
Effective September 15, newly registered domains within Cloudflare will have their Training and Agent crawlers blocked by default on pages featuring advertisements, while Search crawlers will still have access.

The persistence of the user-initiated loophole remains uncertain. Its viability hinges on the assertion that a request to access a page differs fundamentally from a crawler independently sourcing it, a dynamic that now applies across all principal assistants.
Source link: Searchenginejournal.com.





