Research · September 2026
State of AI Readiness
We checked the homepage and robots.txt of 922 well-known sites to see how they treat ChatGPT, Claude and Perplexity: who lets AI crawlers in, who has left a map for them, and who is ready to be cited.
11%
block at least one AI crawler in robots.txt
101 of 922 sites
6%
block a crawler that powers AI search answers
58 sites
46%
publish an /llms.txt
20% also publish /llms-full.txt
77.3
average AI-readiness score, out of 100
median 78 · 7% score 90+
Which AI crawlers get blocked
“Blocked” means the site’s robots.txt fully disallows that crawler. Search and answer crawlers decide whether a site can be surfaced and cited in AI answers; training crawlers only feed the models. 5% of sites block training crawlers while leaving the search ones open.
AI search & answer crawlers
- ChatGPT Search · OAI-SearchBot3% blocked
- ChatGPT (live browsing) · ChatGPT-User5% blocked
- Claude (live browsing) · Claude-User4% blocked
- Perplexity answers · PerplexityBot6% blocked
- Perplexity (live browsing) · Perplexity-User3% blocked
Model-training crawlers
- ChatGPT (training) · GPTBot7% blocked
- Claude (training) · ClaudeBot7% blocked
- Gemini & AI Overviews · Google-Extended7% blocked
- Apple Intelligence · Applebot-Extended6% blocked
- Common Crawl corpus · CCBot8% blocked
Fully blockedSome paths blockedSites that block every AI crawler we check: 1%
Where sites lose points
The AI-readiness checks that well-known sites most often fail or only partly pass. Definitions are on the methodology page.
Most often falls short
- Content dates89% fail or warn
- /llms-full.txt80% fail or warn
- RSS/Atom feed80% fail or warn
- Sitemap freshness59% fail or warn
- Structured data (validated)55% fail or warn
- /llms.txt54% fail or warn
- Heading outline48% fail or warn
- Author / Organization markup48% fail or warn
Most often done right
- Training-crawler stance100% pass
- Consent wall100% pass
- Snippet directives100% pass
- Server-rendered content94% pass
- CDN / firewall gate90% pass
The most and least ready categories
Average AI-readiness score by category. Each links to its leaderboard.
- Marketing & sales84.3
44 sites · 5% block an AI crawler · best: Gong (97)
- Support & messaging APIs84
36 sites · 6% block an AI crawler · best: Sendbird (97)
- Databases & data83.7
43 sites · 0% block an AI crawler · best: Neo4j (97)
- AI & machine learning82.9
48 sites · 2% block an AI crawler · best: Together AI (94)
- Security & identity82.6
45 sites · 0% block an AI crawler · best: Clerk (97)
- Analytics & monitoring81.8
49 sites · 2% block an AI crawler · best: Grafana (94)
- Cloud & hosting79.9
40 sites · 3% block an AI crawler · best: Vercel (94)
- CMS & publishing79.8
36 sites · 6% block an AI crawler · best: Contentstack (94)
- Developer tools78.9
105 sites · 0% block an AI crawler · best: ngrok (94)
- Productivity & collaboration77.9
51 sites · 4% block an AI crawler · best: Shortcut (94)
- Design & no-code76.7
38 sites · 8% block an AI crawler · best: Adalo (92)
- Payments & fintech76.4
50 sites · 4% block an AI crawler · best: FastSpring (97)
- Health, science & public sector74.9
29 sites · 10% block an AI crawler · best: World Bank (89)
- E-commerce & retail73.8
36 sites · 8% block an AI crawler · best: Recharge (92)
- News & media72.7
56 sites · 86% block an AI crawler · best: The Independent (89)
- Travel & marketplaces72
30 sites · 7% block an AI crawler · best: Ashby (92)
- Open source & nonprofits71.5
50 sites · 4% block an AI crawler · best: Wikimedia Foundation (89)
- Education & reference70.6
56 sites · 11% block an AI crawler · best: University of Toronto (89)
- Communication & social69.9
34 sites · 26% block an AI crawler · best: Circle (89)
- Consumer apps & entertainment69.9
46 sites · 26% block an AI crawler · best: The Weather Channel (89)
The most restrictive robots.txt files
Sites that fully disallow the most AI crawlers. Blocking is a legitimate choice (publishers, paywalled and regulated sites often have good reasons); the point is that it should be a decision, not an accident. Each links to that site’s scorecard.
- AllTrailsalltrails.com10 blocked
Consumer apps & entertainment · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- CNNcnn.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- Gentoogentoo.org10 blocked
Open source & nonprofits · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- HuffPosthuffpost.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- NBC Newsnbcnews.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- New Scientistnewscientist.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- The Asahi Shimbunasahi.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- The New York Timesnytimes.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- TikToktiktok.com10 blocked
Communication & social · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- USA Todayusatoday.com10 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- Amazonamazon.com9 blocked
E-commerce & retail · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Perplexity-User, CCBot
- BBCbbc.com9 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- CNBCcnbc.com9 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Google-Extended, PerplexityBot, Applebot-Extended, CCBot
- CNETcnet.com9 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
- DWdw.com9 blocked
News & media · GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Google-Extended, PerplexityBot, Perplexity-User, Applebot-Extended, CCBot
Method and limits
- The sample. These are 922 of the 1,122 recognisable sites in our directory; the other 200 couldn’t be checked (they block automated checks, were unreachable, or told crawlers to stay out). It is a curated list of well-known brands, not a random sample of the web, so read the numbers as “well-known sites”, not “the internet”.
- What we read. Each site’s homepage and robots.txt, fetched by AvocadoScoreBot. Nothing else.
- What “blocks” means. robots.txt fully disallows that crawler (
Disallow: /). Partial rules are shown separately and not counted as blocking. robots.txt is a request that well-behaved crawlers honour; it isn’t enforcement. - Scores. AvocadoScore’s AI-readiness checks (see the methodology). Sites are re-checked about weekly; this page shows data from September 2026.
- Spot a mistake? Email hello@avocadoscore.io, including to have a site removed.
Cite this: AvocadoScore, “State of AI Readiness”, September 2026. https://avocadoscore.io/state-of-ai-readiness
Where does your site land?
The same checks, on your site, free and without signing up. You get a ranked fix list, and you can see how you compare with your category.