Study plan, registered on 2 October 2026. This study has not been run yet. I’m publishing the method before any results exist so it can’t be adjusted to fit them. Results will be added to this page, with the date they were collected.
My October crawler study read what the UK’s ten biggest news sites ask AI crawlers to do. robots.txt is a request, so the next question is whether the servers act on it. This study will check that for each site’s homepage.
What will be measured
On a single date, each of the ten homepages will be requested four times: once with a normal browser user agent, and once each with the published user agents for GPTBot, ClaudeBot and ChatGPT-User. For each request I’ll record the final URL after redirects, the HTTP status code, and whether the response is a block page. A block page means an error status, a challenge or access-denied page, or a response that isn’t the homepage.
Each crawler result will be compared with what that site’s robots.txt said about the same crawler on 2 October 2026. The two agree if robots.txt blocked the crawler and the server refused it, or if robots.txt allowed the crawler and the server served the page. The headline figure will be the number of pairs that disagree.
First pass
On 2 October 2026 all ten homepages loaded in a normal browser. dailymail.co.uk redirected to dailymail.com. The crawler requests haven’t been made yet.
Limits
This covers homepages only, on one date. A block page might come from a service such as Cloudflare rather than from the publisher. All requests will come from one network, so a site that filters by network could treat them differently from requests made elsewhere. The normal browser request from the same network is there as a check on that.
robots.txt is advisory. If a server serves a page to a crawler its robots.txt blocks, that shows a gap between the policy and what the server does. It doesn’t show anything more than that.
What would show the result was wrong
Repeating the same requests to the same homepages on the same day and getting a different number of disagreements.
Data
When the study runs, the full results will be published here as a CSV, one row per site and user agent. The method note covers the sample and the definition of blocked.