Who blocks AI agents?

A scan of the home pages of the top 1,000 websites, fetched as nine different HTTP clients on 3 October 2026.

9%

of sites that serve a browser refuse at least one of curl, node, Python urllib, Python requests or libwww-perl: the libraries agents are built on (45 of 500).

176

sites refuse even a Chrome browser when the request comes from cloud infrastructure, where most agents run.

71.4%

of sites that refuse GPTBot or ClaudeBot say nothing about it in robots.txt (35 of 49).

11.6%

publish an llms.txt (90 of 773 scanned sites).

Why it matters

An AI agent reading a page does not use a browser. It sends a request from a library such as Python's requests or Node's fetch, usually from a cloud server. Many sites treat that request differently from a person's browser: they answer it with a 403, a bot challenge, or a rate-limit, often without the owner knowing. The agent then reports that the page is unavailable, or quietly uses an older copy from somewhere else.

Refusals by client

Of the 500 sites that served our Chrome browser profile, the share that refused each other client with a 401, 403, 407, 429 or 451:

ClientSites refusingShare
curl173.4%
node/undici91.8%
python-urllib265.2%
python-requests428.4%
libwww-perl163.2%
GPTBot408%
ClaudeBot387.6%
Googlebot285.6%

5 sites refused every generic library while serving the browser. Among the top 100 sites that served the browser, 3 of 60 refused at least one library.

Where the generic-library refusals came from: a Cloudflare bot challenge on 13 sites, a Cloudflare edge rule on 0, the origin server on 1, and an unidentified layer (often a CDN or WAF that does not name itself) on 31. A site can appear under more than one.

Examples of sites that serve a browser but refuse at least one library: wikipedia.org, endpoints.news, openai.com, snapchat.com, criteo.com, hosting24.com, medium.com, vungle.com, wikimedia.org, imdb.com, wiley.com, ip-api.com.

Silent AI-crawler blocks

49 sites (9.8%) refused GPTBot or ClaudeBot. For 35 of them, robots.txt does not disallow that crawler, so a well-behaved crawler that reads robots.txt is still turned away at the door. Examples: wikipedia.org, endpoints.news, openai.com, wordpress.com, hosting24.com, rubiconproject.com, wikimedia.org, wiley.com, weebly.com, un.org, 163.com, wp.com.

Some of these will be sites that verify a crawler's IP address and refuse the User-Agent when it arrives from elsewhere, as ours did. That is a reasonable policy against impersonation, but it means agents that browse on a user's behalf with those names are refused too.

What robots.txt says

124 of 773 sites (16%) disallow at least one AI crawler from the whole site in robots.txt:

CrawlerSites disallowing it
CCBot107
GPTBot98
ClaudeBot94
Google-Extended87
Applebot-Extended78
PerplexityBot75
anthropic-ai73
ChatGPT-User66
Claude-User62

llms.txt adoption

90 sites (11.6%) serve a real llms.txt rather than an HTML page at that path, including cloudflare.com, github.com, fastly.net, wordpress.org, digicert.com, adobe.com, opera.com, sentry.io, wordpress.com, dropbox.com, kaspersky.com, gravatar.com.

Redirects that drop the request body

Of 621 sites that redirect plain http to https, 598 (96.3%) use a 301, 302 or 303. Clients turn a POST into a GET when following those, so an agent that posts to an http:// URL loses its body. A 307 or 308 keeps it.

Cloud-origin blocks

176 sites refused a full Chrome User-Agent (13 of them with a 429 rate limit rather than a block), which means they judge the request by where it comes from rather than what it says it is. Agents running in a cloud provider meet the same wall. Examples: google.com, microsoft.com, appsflyersdk.com, tiktok.com, cloudflare.net, googlesyndication.com, chatgpt.com, samsung.com, windows.com, b-cdn.net, office365.com, okcdn.ru. A further 97 sites timed out or returned a server error to the browser profile and are left out of the percentages above.

Method

Check your own site

The same nine-client scan is available for any page. The free scan covers one page; the report covers up to 20 pages with causes and fixes; monitoring re-checks daily and tells you when something changes.

Free scan →   $69 report · Monitoring, $19/month