On 2 October 2026 the Telegraph’s robots.txt blocked 14 of the 15 AI crawlers I checked. The fifteenth was Google-Extended, the token Google uses to let sites opt out of having their content used for its AI models. The Telegraph didn’t simply leave Google-Extended out of its file. It added a rule that explicitly allows it. Someone made that choice, so I want to look at it as one.
The choice
Put plainly, the file says that no AI company other than Google may use the Telegraph’s site for training, and Google may. It’s the only site in the sample with an explicit allow for Google-Extended.
It isn’t the only site that lets Google-Extended through. Only 5 of the 10 sites block it, which makes it the least-blocked training crawler in the study. ClaudeBot and CCBot, by comparison, are each blocked by 9 of 10. What’s different about the Telegraph is that it wrote the exception down. The other sites that don’t block Google-Extended have no rule about it either way.
What I’m leaving out
It’s tempting to guess at a reason, such as a commercial arrangement or a view about Google’s importance to UK news. I’m not going to, because robots.txt records the rule and says nothing about why it was written. All the file shows is that the exception exists and that it was made explicit.
What a later snapshot would need to show
I don’t think “clever” can be judged from one file. Before I’d use that word, a later snapshot would need to show the Google-Extended allow still in place, along with the blocks on the other crawlers. Even then, the file would only show that the policy held. Whether it paid off would need evidence that robots.txt can’t give, since the file records requests and says nothing about results. On robots.txt alone, the most I can say is that the exception is deliberate and still there.
This post isn’t a prediction. Dated calls go on the predictions page, where they get marked.
The file, the sample and the definition of blocked are in the study: Which AI crawlers do the UK’s biggest news sites block?