Most large AI companies run several crawlers, and each one does a different job. A lot of confusion about “blocking AI” comes from people meaning different crawlers without saying so. OpenAI is a useful example because it runs three.
GPTBot collects pages that may be used to train OpenAI’s models. Blocking it means a site is asking for its pages to stay out of future training data.
OAI-SearchBot builds the index that ChatGPT’s search features draw on. Blocking it asks for a site’s pages to be left out of that index.
ChatGPT-User fetches a page when a ChatGPT user’s request needs it. Blocking it asks ChatGPT not to read the page on that user’s behalf.
A publisher can block any combination of the three. So “we block OpenAI” doesn’t tell you much until you know which crawler it means.
What the biggest UK news sites did
On 2 October 2026 I read the robots.txt files of the UK’s ten biggest newsbrands. GPTBot was blocked by 7 of the 10, OAI-SearchBot by 6 and ChatGPT-User by 4. Each publisher wrote one file, but most treated the three crawlers differently, and the training crawler was the one blocked most often.
Anthropic’s crawlers showed the same split. ClaudeBot, the training crawler, was blocked by 9 of 10. Claude-SearchBot was blocked by 5 and Claude-User by 4.
Why it matters
The numbers suggest these publishers are more concerned about their journalism being used for training than about an assistant reading a page when a reader asks for it. You may or may not agree with that position, but you can only see it if you look at each crawler separately.
If you’re reporting that a publisher “blocks AI”, it helps to say which crawler. If you’re reading a count of blockers, check whether it names the crawlers it counted.
The full table and definitions are in the study: Which AI crawlers do the UK’s biggest news sites block?