Skip to main content

AI Crawler Access Index

Do law firms and dental practices let AI assistants read their websites? We checked robots.txt and llms.txt on the 625 websites that showed up when we asked ChatGPT and Google for the best local practices. Collected 2026-09-16.

Robots.txt rules checked for AI crawler user agents

Key findings

  • 2.1%

    of 574 practice websites block at least one AI crawler in robots.txt

  • 32.6%

    of the 43 directories and multi-city brands in the same results do

  • 35.9%

    of practice websites serve an llms.txt file

  • 17.8%

    of practice websites have an llms.txt that names an SEO plugin as its author

  • Law firms were more likely than dental practices to publish llms.txt: 44.1% against 27.8%.
  • Plugin-generated files by author: Yoast 62, Rank Math 23, AIOSEO 15.
  • 11.7% of practice websites had no usable robots.txt at all, which leaves every crawler allowed by default.
  • Only 4.7% of practice websites declared Cloudflare-style content signals (for example ai-train=no).

Blocked in robots.txt, by crawler

Share of websites whose robots.txt blocks the whole site for each crawler, either by name or through a rule for all bots.

Percentage of websites blocking each AI crawler
CrawlerOperatorPurposeLaw firmsDentalDirectories
GPTBotOpenAItraining0.7%1.7%20.9%
OAI-SearchBotOpenAIsearch0.0%0.0%9.3%
ChatGPT-UserOpenAIuser fetch0.0%0.0%11.6%
ClaudeBotAnthropictraining0.3%2.4%18.6%
Claude-SearchBotAnthropicsearch0.0%0.0%9.3%
PerplexityBotPerplexitysearch0.0%1.7%16.3%
Google-ExtendedGoogletraining0.3%0.0%11.6%
Applebot-ExtendedAppletraining0.3%2.1%16.3%
CCBotCommon Crawltraining0.7%1.7%23.3%
BytespiderByteDancetraining0.7%2.1%23.3%
meta-externalagentMetatraining0.7%2.1%11.6%
AmazonbotAmazontraining0.7%0.3%9.3%

What to do with this

Most practices are not the problem. Almost none block AI crawlers, so if your practice is missing from AI answers, robots.txt is unlikely to be why. Our ChatGPT vs Google study found that directories and publications shape those answers more than practice websites do.

Check an llms.txt your plugin wrote. An auto-generated file often lists a sitemap and little else. If you keep one, make it describe your services, locations and the pages that answer patients' or clients' real questions.

Decide bot by bot. Search bots and training bots do different jobs. The free AI search visibility checker walks through the other factors that decide whether assistants cite you.

Methodology

  • Sample: every website domain that appeared in ChatGPT's answers or Google's map pack and top results for “best dentist” and “best personal injury lawyer” in 20 US metros, collected for our ChatGPT vs Google study.
  • 625 domains checked on 2026-09-16; 8 could not be reached and are excluded. Each site received one request for /robots.txt and one for /llms.txt, with an identifying user agent.
  • Directories and multi-city brands: domains on our published directory list, or appearing in 3 or more metros. Everything else counts as a practice website.
  • A crawler counts as blocked when the robots.txt group that applies to it (its own group, or the * group if it has none) disallows the whole site and no Allow rule reopens it, following RFC 9309. A missing robots.txt or a 4xx response counts as allowed.
  • llms.txt counts when /llms.txt returns a 200 plain-text response that is not an HTML error page.
  • Limits: firewall and CDN-level blocking is invisible to this method, and a site can change its rules at any time.

Frequently asked questions

Cite this research

Space Agency. “AI Crawler Access Index: Law Firm and Dental Websites.” Published 2026-09-16. https://www.spaceagency.online/research/ai-crawler-access/

Journalists and researchers may use any figure on this page with a link to this URL. The underlying data is available on request.