Study plan, registered on 2 October 2026. This study has not been run yet. I’m publishing the method before any results exist so it can’t be adjusted to fit them. Results will be added to this page, with the date they were collected.
Nine of the UK’s ten biggest newsbrands blocked at least one AI crawler on 2 October 2026, according to my October crawler study. This study will compare those ten with ten high-reach UK sites that fail a credibility test, to see whether the two groups block AI crawlers at different rates.
The credibility test
This test is set before any of the comparison sites are chosen and before any of their robots.txt files are read. A site fails if it meets at least two of these conditions:
- It has no named editor or masthead page that can be reached from the homepage.
- It has no corrections policy or corrections page.
- At least two of its stories have been rated false or misleading by Full Fact in the last 24 months, with each rating linked.
How the sites will be chosen
The comparison sites will be taken from a named, published ranking of UK website reach, which I’ll state when the list is fixed. Each site’s result against the test will be recorded with links to the evidence. Once the list of ten is fixed, it won’t change. All 20 robots.txt files will then be read on a single date and scored using the same definition of blocked as the first study, set out in the method note.
What will be published
The block rates for the two groups, side by side, together with the evidence for each comparison site. The study won’t discuss why either group made the choices it did.
Limits
The test looks at whether a site has a named editor, a corrections page and a fact-check record. It isn’t a judgement about a site’s politics. Block rates show what robots.txt files ask for, not what crawlers do. Ten sites in each group is enough for a comparison but not for cause and effect.
What would show the result was wrong
Reading the same 20 files on the same date and getting different block rates for either group.