This is the standing method note for the AI crawler series. Each study links here instead of restating the rules, and the rules don’t widen between studies.
The sample
The ten biggest UK newsbrands by audience in August 2026, as listed in Press Gazette’s monthly ranking using Ipsos iris data: news.sky.com, theguardian.com, dailymail.com, independent.co.uk, bbc.co.uk, mirror.co.uk, uk.news.yahoo.com, thesun.co.uk, express.co.uk and telegraph.co.uk. The sample is frozen: re-runs check these same ten sites even if the ranking changes, so a change in the numbers means a publisher changed its file, not that the sample changed under it.
The snapshot
Each study reads every robots.txt file on a single stated date. Results are a snapshot of that day, because these files change.
What “blocked” means
A crawler counts as blocked when the rules that apply to it disallow the whole site (Disallow: /), either in a group that names it or through the catch-all group when it isn’t named. An explicit Allow: / for that crawler overrides. Anything short of a whole-site disallow counts as not blocked, even when some paths are restricted.
The crawlers
This series checks 15 crawlers: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, anthropic-ai, Google-Extended, PerplexityBot, Perplexity-User, CCBot, Applebot-Extended, Bytespider, Meta-ExternalAgent and Amazonbot.
What robots.txt does not show
- robots.txt is a request, not a wall. It shows what a publisher asks for, not what crawlers do.
- Network-level blocks don’t show up in it. A site can allow a crawler in robots.txt and still turn it away at the door.
- It says nothing about whether any crawler actually obeys.
What this site will not claim
- Whether any crawler respects these files.
- Whether blocking or allowing a crawler caused any change in a site’s traffic or citations.
- Anything about sites or crawlers outside the stated sample.
- Any figure that doesn’t trace back to a published dataset.
Sources and tools
The sources and tools behind each study are listed on the What I use page.
Raw data
Every study in the series publishes its full data as a CSV download, so anyone can check the findings or re-run the comparison.