robots.txt is a request

My October study counted what the UK’s ten biggest news sites ask AI crawlers to do. Its limits section made one point I think deserves a post of its own: robots.txt is a request. It isn’t a wall.

What the file shows

A robots.txt file sets out a site’s policy for each crawler on the day you read it. It’s precise about intent. You can see which companies are named, which crawlers are refused, and whether a refusal covers the whole site or only some paths. As evidence of what a publisher wants, it’s hard to beat, because the publisher wrote it and anyone can read it at a known address.

What it doesn’t show

The file says nothing about enforcement. A site can block crawlers at the network level, and none of that appears in robots.txt. A site can also disallow everything in the file while its server still responds to any crawler that ignores the file. On 2 October 2026 the Daily Mail’s file blocked all 15 crawlers I checked. That was the strictest file in the sample, but I didn’t test whether anything enforces it, so the study can’t say.

The file also can’t tell you whether crawlers follow it. Finding that out would need server logs or test requests, which is a different kind of study.

Why a later citation wouldn’t contradict the file

Suppose a site blocks a training crawler today, and later some of its reporting seems to turn up in an AI product. That wouldn’t prove the file was ignored. robots.txt only covers crawling of that site from the time the rule is in place. It doesn’t cover text collected before the rule existed, or copies of the same text published somewhere else. I’m not saying any blocked site is being cited anywhere. My point is that if it happened, the file on its own couldn’t explain how.

That’s why claims in my research come with a date, a definition of blocked, and the reminder that the file is a request. Without those, it’s easy to claim more than robots.txt can support.

The full limits, sample and definitions are in the study, Which AI crawlers do the UK’s biggest news sites block? and in the method note.