In my October study, a crawler counts as blocked when the rules that apply to it disallow the whole site. That means Disallow: / in a group that names the crawler, or in the catch-all group if nothing names it. This post explains why I think a definition short enough to fit in one sentence is the right one to use.
The main reason is that anyone can check it. You open the robots.txt file, find the group that applies to the crawler, and see whether the whole site is disallowed. There’s no judgement involved, so two people who disagree about AI will still get the same count. If the definition leaves room for argument, so do the numbers built on it.
What “we block AI” can hide
When a publisher says it blocks AI, the claim sometimes rests on something looser than a full-site disallow. I’ve seen three versions of this.
The first is a path rule. Keeping a crawler out of /search/ or a few directories still leaves the homepage and every article open to it. If that counted as a block, almost every site would block almost every crawler, and the word would stop meaning much.
The second is naming a crawler without disallowing the whole site. A line that says User-agent: GPTBot looks firm in a screenshot. If the rules under it stop short of Disallow: /, GPTBot can still crawl most of the site.
The third is a block on the server rather than in the file. Some sites turn bots away at the network level. That is a real block, but it doesn’t appear in robots.txt, so it can’t be counted from robots.txt. Putting both kinds in one number makes the number impossible to check.
It can also go the other way. A file with no AI rules at all can still block every AI crawler if its catch-all group disallows everything it doesn’t name. That’s why the definition looks at the rules that apply to each crawler and ignores which crawlers the file happens to mention.
An example from the study
On 2 October 2026, the Sun’s robots.txt blocked everything by default and then let in a list of named crawlers. OpenAI’s three crawlers were on that list, with some paths restricted. Every other AI crawler I checked fell under the default block. With a loose definition you could describe that file as “the Sun blocks AI” or “the Sun welcomes AI”, and both would be partly true. Checking each crawler against the one-sentence definition gives a clear answer for each of them.
So when you see a count of which sites block which crawlers, look for the definition first. If you can’t find one, treat the count with caution.
The study this comes from is Which AI crawlers do the UK’s biggest news sites block? The definition is also set out in the method note, so later studies use the same one.