A scan of the home pages of the top 1,000 websites, fetched as nine different HTTP clients on 3 October 2026.
of sites that serve a browser refuse at least one of curl, node, Python urllib, Python requests or libwww-perl: the libraries agents are built on (45 of 500).
sites refuse even a Chrome browser when the request comes from cloud infrastructure, where most agents run.
of sites that refuse GPTBot or ClaudeBot say nothing about it in robots.txt (35 of 49).
publish an llms.txt (90 of 773 scanned sites).
An AI agent reading a page does not use a browser. It sends a request from a library such as Python's requests or Node's fetch, usually from a cloud server. Many sites treat that request differently from a person's browser: they answer it with a 403, a bot challenge, or a rate-limit, often without the owner knowing. The agent then reports that the page is unavailable, or quietly uses an older copy from somewhere else.
Of the 500 sites that served our Chrome browser profile, the share that refused each other client with a 401, 403, 407, 429 or 451:
| Client | Sites refusing | Share |
|---|---|---|
| curl | 17 | 3.4% |
| node/undici | 9 | 1.8% |
| python-urllib | 26 | 5.2% |
| python-requests | 42 | 8.4% |
| libwww-perl | 16 | 3.2% |
| GPTBot | 40 | 8% |
| ClaudeBot | 38 | 7.6% |
| Googlebot | 28 | 5.6% |
5 sites refused every generic library while serving the browser. Among the top 100 sites that served the browser, 3 of 60 refused at least one library.
Where the generic-library refusals came from: a Cloudflare bot challenge on 13 sites, a Cloudflare edge rule on 0, the origin server on 1, and an unidentified layer (often a CDN or WAF that does not name itself) on 31. A site can appear under more than one.
Examples of sites that serve a browser but refuse at least one library: wikipedia.org, endpoints.news, openai.com, snapchat.com, criteo.com, hosting24.com, medium.com, vungle.com, wikimedia.org, imdb.com, wiley.com, ip-api.com.
49 sites (9.8%) refused GPTBot or ClaudeBot. For 35 of them, robots.txt does not disallow that crawler, so a well-behaved crawler that reads robots.txt is still turned away at the door. Examples: wikipedia.org, endpoints.news, openai.com, wordpress.com, hosting24.com, rubiconproject.com, wikimedia.org, wiley.com, weebly.com, un.org, 163.com, wp.com.
Some of these will be sites that verify a crawler's IP address and refuse the User-Agent when it arrives from elsewhere, as ours did. That is a reasonable policy against impersonation, but it means agents that browse on a user's behalf with those names are refused too.
124 of 773 sites (16%) disallow at least one AI crawler from the whole site in robots.txt:
| Crawler | Sites disallowing it |
|---|---|
| CCBot | 107 |
| GPTBot | 98 |
| ClaudeBot | 94 |
| Google-Extended | 87 |
| Applebot-Extended | 78 |
| PerplexityBot | 75 |
| anthropic-ai | 73 |
| ChatGPT-User | 66 |
| Claude-User | 62 |
90 sites (11.6%) serve a real llms.txt rather than an HTML page at that path, including cloudflare.com, github.com, fastly.net, wordpress.org, digicert.com, adobe.com, opera.com, sentry.io, wordpress.com, dropbox.com, kaspersky.com, gravatar.com.
Of 621 sites that redirect plain http to https, 598 (96.3%) use a 301, 302 or 303. Clients turn a POST into a GET when following those, so an agent that posts to an http:// URL loses its body. A 307 or 308 keeps it.
176 sites refused a full Chrome User-Agent (13 of them with a 429 rate limit rather than a block), which means they judge the request by where it comes from rather than what it says it is. Agents running in a cloud provider meet the same wall. Examples: google.com, microsoft.com, appsflyersdk.com, tiktok.com, cloudflare.net, googlesyndication.com, chatgpt.com, samsung.com, windows.com, b-cdn.net, office365.com, okcdn.ru. A further 97 sites timed out or returned a server error to the browser profile and are left out of the percentages above.
https://<domain>/ per client, following up to 5 redirects, with a 10-second timeout. Clients differ only in their User-Agent header; every request also carries an X-Scanner header linking to this page.* group if none does, disallows /. Rules for specific paths are not counted.The same nine-client scan is available for any page. The free scan covers one page; the report covers up to 20 pages with causes and fixes; monitoring re-checks daily and tells you when something changes.