Published 7 October 2026.
When people argue about AI crawlers, they picture someone making a call. Publishers do. On 2 October 2026, nine of the UK’s ten biggest newsbrands blocked at least one AI crawler in robots.txt (my study). Those files name bots, split training from retrieval and carve out exceptions. Someone sat down and decided.
A plumber’s website isn’t like that. My bet is most small business sites have no AI lines in robots.txt at all. Under the robots.txt standard, a crawler with no matching rule is allowed (RFC 9309). Silence is a yes.
I call this the default vote: a site that says nothing about AI has voted for whatever its platform shipped.
So the platforms decide
- Squarespace has a setting called “Block known artificial intelligence crawlers”. It asks bots such as GPTBot, ClaudeBot and CCBot to stay out. It’s off by default (Squarespace help, read 7 October 2026).
- Cloudflare went the other way. In July 2025 it said it would block AI bots by default on sites it serves (MIT Technology Review).
Two owners who care equally little can end up with opposite answers, depending on which company sits in front of their site.
What it means if you measure this
A count of small business sites that “allow” AI crawlers is mostly a count of platform defaults. It isn’t evidence that owners want to be in AI answers. Read those numbers as a map of Squarespace, Wix, Shopify and Cloudflare settings, not of intent.
The call
- Claim: Squarespace will still ship new sites with “Block known artificial intelligence crawlers” switched off.
- Called: 7 October 2026. Check date: 7 October 2027.
- Test: Read Squarespace’s help page for the setting on the check date.
- Right if: the page says the setting is off by default.
- Wrong if: it’s on by default, or new Squarespace sites block AI crawlers some other way.
- Baseline: off by default on 7 October 2026.
I’ll mark this on the check date and leave the result up either way.
Added 10 October 2026: this call is P6 on the scoreboard.