Which UK news sites block GPTBot and other AI crawlers?

Checked 2 October 2026.

Nine of the UK’s ten biggest newsbrands block at least one AI crawler. But most of them block the bots that train AI models far more often than the bots that fetch pages when someone asks ChatGPT or Claude a question.

I checked the robots.txt file of each of the UK’s ten biggest newsbrands on 2 October 2026 against 15 AI crawlers. Here’s what they’re actually blocking.

The headline numbers

  • 9 of 10 block at least one AI crawler. The Independent is the exception: its robots.txt has no AI-specific rules at all.
  • The Daily Mail blocks all 15 crawlers I checked.
  • The Telegraph blocks 14 of the 15 AI crawlers I checked but explicitly allows Google-Extended, the token that controls whether Google can use its content for Gemini.
  • Only 5 of 10 block Google-Extended. It’s the least-blocked training token in the sample.
  • ClaudeBot, PerplexityBot and CCBot are each blocked by 9 of 10.

Publishers block training more than answers

Each AI company runs more than one crawler. One collects pages to train models. Another builds a search index. A third fetches a page when a user asks about it. UK publishers treat them very differently.

CompanyTraining crawlerSearch crawlerUser-triggered fetcher
OpenAIGPTBot: 7 of 10OAI-SearchBot: 6 of 10ChatGPT-User: 4 of 10
AnthropicClaudeBot: 9 of 10Claude-SearchBot: 5 of 10Claude-User: 4 of 10
PerplexityNot in studyPerplexityBot: 9 of 10Perplexity-User: 4 of 10
GoogleGoogle-Extended: 5 of 10Not in studyNot in study
Number of the ten sites that block each crawler from the whole site.

Four sites (Sky News, the BBC, the Mirror and the Daily Express) block ClaudeBot but leave Claude-SearchBot open. Only Sky News does the same with OpenAI: it blocks GPTBot but not OAI-SearchBot.

The pattern is the same across all three companies. Publishers are keener to stop their journalism training models than to stop it being fetched and cited in answers. That makes sense if the goal is to keep the referral traffic that AI answers can send while refusing free training data.

Two outliers

The Sun blocks almost everything by default and lets in a named list of crawlers. OpenAI’s three crawlers are on that list, with some paths restricted. Every other AI crawler in the sample is blocked.

The Telegraph blocks every other AI crawler I checked, but adds a rule that specifically lets Google-Extended in. No other site in the sample does that.

All ten sites

SiteAI crawlers blocked (of 15)GPTBotClaudeBotPerplexityBotGoogle-Extended
Daily Mail15YesYesYesYes
The Telegraph14YesYesYesNo (explicitly allowed)
BBC13YesYesYesYes
The Sun12NoYesYesYes
Sky News10YesYesYesYes
The Guardian10NoYesYesNo
Yahoo! (UK news)10YesYesYesYes
Mirror8YesYesYesNo
Daily Express7YesYesYesNo
The Independent0NoNoNoNo

The full data for all 15 crawlers is available as a CSV download.

Method

  • The sample is the ten biggest UK newsbrands by audience in August 2026, as listed in Press Gazette’s monthly ranking using Ipsos iris data.
  • Each robots.txt file was read on 2 October 2026.
  • The files checked were at news.sky.com, theguardian.com, dailymail.com, independent.co.uk, bbc.co.uk, mirror.co.uk, uk.news.yahoo.com, thesun.co.uk, express.co.uk and telegraph.co.uk.
  • The 15 crawlers were GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, anthropic-ai, Google-Extended, PerplexityBot, Perplexity-User, CCBot, Applebot-Extended, Bytespider, Meta-ExternalAgent and Amazonbot.
  • A crawler counts as blocked when the rules that apply to it disallow the whole site (Disallow: /), either in a group that names it or through the catch-all group when it isn’t named.

The standing method note for this series covers the sample, the definitions and the limits in one place; the definitions above follow it.

Limits

  • robots.txt is a request, not a wall. It shows what a publisher asks for, not what crawlers do.
  • Some sites block crawlers at the network level too. That doesn’t show up here.
  • Ten sites is a small sample. It covers the biggest audiences, not the whole UK news market.
  • These files change. This is a snapshot from one day.

What happens next

I’ll re-run this check on the 2nd of every month against the same ten sites and the same 15 crawlers, and keep every month’s results. One of my public predictions rests on it.

Follow along on LinkedIn

I share new research, predictions and what I’m seeing in AI search on LinkedIn first.