In my October study, nine of the ten robots.txt files had the same basic structure. Crawlers were allowed in by default, and the file named the ones the site wanted to keep out. The Sun’s file, which I read on 2 October 2026, worked the other way round.
Blocked by default
The Sun’s catch-all rule blocks every crawler the file doesn’t name. The file then lists named crawlers that are allowed in. A crawler that isn’t on the list is blocked, even though no rule mentions it.
Which AI crawlers are allowed
Of the AI crawlers in my study, only OpenAI’s are on the Sun’s list. GPTBot, OAI-SearchBot and ChatGPT-User are all allowed, with some paths restricted. Every other AI crawler I checked, including Anthropic’s, Perplexity’s and Common Crawl’s, is blocked by the default rule.
How this kind of file behaves
A file built this way differs from the others in two ways. First, every crawler that gets in has been added to the list by name. In a file that allows everything by default, you can’t tell whether a crawler is allowed on purpose or just hasn’t been noticed. In the Sun’s file, each allowed crawler was written in.
Second, new crawlers start out blocked. If a new AI company launches a crawler, the Sun’s file already blocks it. At sites that allow by default, a new crawler can crawl until someone adds a rule for it.
What stands out
Most files in the sample read like general policies, for example blocking training crawlers while allowing fetchers. The Sun’s file treats one company differently from all the others. My data shows how the file is set up but not why, so I’ll stop there. This is one site on one date, and I’m not drawing any wider conclusion from it.
The sample, the definitions and the other nine files are in the study: Which AI crawlers do the UK’s biggest news sites block?