Some are, but fewer than the headline number suggests. We read the robots.txt file on 6,497 US dental practice websites in July 2026. 11.4% block at least one AI crawler outright. Narrow that to the crawlers run by OpenAI, Anthropic and Perplexity and it drops to 7.6%, about one site in thirteen. Another 11.4% publish no robots.txt at all, so their access rules are undeclared rather than closed.
A robots.txt file is a short text file at the root of a website that tells crawlers which parts they may read. OpenAI, Anthropic, Perplexity and Google each publish the names of their crawlers and what each one is for, which makes that one file a real gate: a crawler that respects the file and is turned away cannot read the page, and an engine with nothing to read has nothing to quote. Why a practice stays out of AI answers is a longer question, and we answer it separately in why isn't my dental practice showing up in ChatGPT.
This page is about measurement. On 22 July 2026 we requested /robots.txt from 7,632 US dental practice domains in our database, read the file wherever a server returned one, and applied the rules set out in RFC 9309, the robots.txt standard. 1,135 domains could not be read at all, through 401 and 403 responses, server errors, DNS failures or timeouts. We dropped those from the sample instead of counting them as open, which leaves n = 6,497. The wider method behind our data is in how we measure AI visibility.
How many dental websites block AI crawlers?
739 of the 6,497 sites we could read, 11.4%, block at least one of the crawlers we tested at the root of the site. All but one of those 739 block an AI crawler specifically. Another 739, also 11.4%, return no robots.txt file at all. The largest group by far, 45.4%, blocks a few paths and leaves the site itself open to every crawler we tested.
| What the robots.txt declares | Sites | Share of 6,497 |
|---|---|---|
Some paths blocked, site itself open (Disallow: /wp-admin/ and similar) | 2,950 | 45.4% |
| No restriction at all on any tested crawler | 2,069 | 31.8% |
| At least one crawler blocked at the root | 739 | 11.4% |
| No robots.txt file (404 or a page served instead) | 739 | 11.4% |
The 45.4% in the top row is the least alarming line on this page. We went back and reread the raw files for 2,889 of those 2,950 sites, the ones still held in our fetch store: 1,546 of them, 53.5%, disallow a WordPress admin path, and the exact rule Disallow: /wp-admin/ appears on 1,456. We counted those separately from a root block, because reporting routine housekeeping as a practice shutting the door would turn an 11.4% finding into a 56.8% one.
Which AI crawlers are dental sites blocking?
Mostly not the ones that answer patient questions. The most blocked crawler in our sample is Bytespider, run by ByteDance, at 11.1%. OAI-SearchBot, the crawler OpenAI uses to surface websites in ChatGPT's search results, is blocked by 0.4% of sites. PerplexityBot, which Perplexity documents as the crawler that surfaces and links websites in its results, is blocked by 4.2%.
| Crawler | Operated by | What it does | Sites blocking it |
|---|---|---|---|
| Bytespider | ByteDance | Not on the published crawler list of ChatGPT, Claude or Perplexity | 722 (11.1%) |
| ClaudeBot | Anthropic | Collects web content that may contribute to training Claude | 484 (7.4%) |
| meta-externalagent | Meta | Not on the published crawler list of ChatGPT, Claude or Perplexity | 483 (7.4%) |
| CCBot | Common Crawl | Not on the published crawler list of ChatGPT, Claude or Perplexity | 471 (7.2%) |
| GPTBot | OpenAI | Crawls content that may be used to train OpenAI's models | 458 (7.0%) |
| anthropic-ai | Anthropic | Older token, not in Anthropic's current crawler documentation | 290 (4.5%) |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity results, and is not used for training data | 272 (4.2%) |
| Applebot-Extended | Apple | Not on the published crawler list of ChatGPT, Claude or Perplexity | 239 (3.7%) |
| Google-Extended | Gemini training and grounding in Gemini apps; Google says it does not affect Google Search | 223 (3.4%) | |
| ChatGPT-User | OpenAI | Fetches a page when a user asks ChatGPT to look at it, rather than crawling automatically | 29 (0.4%) |
| Perplexity-User | Perplexity | Fetches a page in response to a user's question | 27 (0.4%) |
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT's search results | 26 (0.4%) |
| Bingbot | Microsoft | Not on the published crawler list of ChatGPT, Claude or Perplexity; it crawls for Bing's search index | 21 (0.3%) |
Read down that table and the shape of the blocking becomes clear. The five most blocked crawlers are Bytespider, ClaudeBot, meta-externalagent, CCBot and GPTBot. Only two of those five belong to ChatGPT, Claude or Perplexity, and both of the two are training crawlers. The ones that decide whether ChatGPT or Perplexity can put a practice into an answer sit near the bottom of the table.
Taken together, 496 sites, 7.6%, block at least one crawler run by OpenAI, Anthropic or Perplexity. The other 243 blocking sites turn away only crawlers from other companies. Two patterns account for most of the blocking: 240 sites run the same seven-name blocklist of Bytespider, CCBot, ClaudeBot, GPTBot, PerplexityBot, anthropic-ai and meta-externalagent, and 237 block Bytespider on its own.
Crawler roles per OpenAI, Anthropic, Perplexity, Google; operator names for Bytespider, CCBot, meta-externalagent and Applebot per Cloudflare's bot reference. Blocking counts: Citevio study, July 2026.
Does blocking GPTBot remove my practice from ChatGPT?
No. OpenAI runs three crawlers with different jobs, and GPTBot is the training one. In our sample 458 dental sites block GPTBot, and 432 of them leave OAI-SearchBot open. Only 26 sites block both. For most of those sites the rule in place is an opt-out from model training, not a block on ChatGPT's answers.
| OpenAI crawler | What OpenAI says it does | Dental sites blocking it |
|---|---|---|
| GPTBot | Crawls content that may be used to train OpenAI's models | 458 (7.0%) |
| ChatGPT-User | Fetches a page when a user asks ChatGPT to visit it, rather than crawling automatically | 29 (0.4%) |
| OAI-SearchBot | Surfaces websites in ChatGPT's search results | 26 (0.4%) |
Perplexity splits its crawlers the other way around. Its documentation describes PerplexityBot as the crawler that surfaces and links websites in Perplexity results, and says it is not used to collect training data. That makes a PerplexityBot block the most consequential rule in the whole table, and it is also the most common of the three assistant blocks: 272 dental sites, 4.2%, turn it away. The 27 sites blocking Perplexity-User may get less than they bargained for, because Perplexity says that fetcher generally ignores robots.txt on the grounds that a person, not a crawler, asked for the page.
Anthropic separates its crawlers the same way, with ClaudeBot for training and Claude-SearchBot for its search index. Our July run tested ClaudeBot and the older anthropic-ai token, not Claude-SearchBot, so read the Anthropic rows as a measure of training opt-outs and not of search access.
What does having no robots.txt mean for AI access?
It means nothing has been declared. 739 dental sites, 11.4% of our sample, returned no robots.txt file. That is not a block. RFC 9309 puts it plainly: "If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server." The practical loss is control, not access.
We report that group separately instead of folding it into the open column, because the two states are not the same thing for a practice owner. An open file is a written decision. A missing file records no decision at all, and it leaves nothing to check against the day a plugin update, a new host or an agency writes a robots.txt on your behalf.
Source: RFC 9309, Robots Exclusion Protocol, 2022.
Does AI crawler blocking vary by state?
Yes, and the spread is wide. Of the four states where we could read at least 300 sites, California blocks at the highest rate: 15.0% of 1,751 sites, against 7.0% of 1,151 in Texas. On the narrower measure, crawlers run by OpenAI, Anthropic or Perplexity, California sits at 11.6% and Texas at 4.8%.
| State | Sites read | Blocks at least one of the 13 crawlers | Blocks an OpenAI, Anthropic or Perplexity crawler | No robots.txt |
|---|---|---|---|---|
| California | 1,751 | 263 (15.0%) | 203 (11.6%) | 222 (12.7%) |
| Illinois | 300 | 36 (12.0%) | 17 (5.7%) | 41 (13.7%) |
| Arizona | 474 | 40 (8.4%) | 30 (6.3%) | 50 (10.5%) |
| Texas | 1,151 | 80 (7.0%) | 55 (4.8%) | 139 (12.1%) |
Two things stand out in that table. The gap between the two blocking columns is widest in Illinois, where 19 of the 36 blocking sites turn away only crawlers that no assistant uses to answer a patient. And the "no robots.txt" column barely moves between states, from 10.5% to 13.7%, while the blocking column runs from 7.0% to 15.0%. In our reading that fits what the missing file usually is: not a decision anyone made, but a site nobody configured. Blocking is also a little more common among practices with fewer reviews, at 13.0% in the quarter with 25 reviews or fewer against 8.7% among practices with more than 315.
Is a crawler block the reason my practice is not in ChatGPT?
For most practices, no. Nearly nine in ten of the sites we read, 5,758 of 6,497, block nothing at the root. Being readable is where visibility starts, not where it lands: SOCi's 2026 Local Visibility Index found that 1.2% of local business locations get recommended on ChatGPT and 7.4% on Perplexity. Far more practices are readable than are recommended.
Our own city scans point the same way. In June 2026 we scanned 212 dental practices in Charlotte, NC and Columbus, OH, and 6 of them blocked AI crawlers. That scan checked a shorter crawler list with a simpler reader and covered two metros, so treat it as a city snapshot beside the national picture on this page. The Charlotte result is the one worth keeping: 147 practices had a site fully open to AI crawlers, and ChatGPT named 2 of them.
Sources: SOCi 2026 Local Visibility Index (~350,000 locations, 2,751 brands), reported by Search Engine Land, January 2026. SOCi measured multi-location brands across industries, not dental practices, so read it as the local-business baseline rather than a dental figure. City numbers: Charlotte and Columbus reports, Citevio scan, June 2026.
robots.txt is a declaration, not enforcement. A server, a firewall or a bot-management rule can still refuse a crawler that the file welcomes, and this method would never know. A separate study we ran found that even open robots.txt files don't guarantee access: what happens when the file says yes and the server says no.
How do I read my own robots.txt in one minute?
Open yoursite.com/robots.txt in a browser. Look for lines that start with User-agent:, then read the rules underneath each one. A line reading Disallow: / blocks that crawler from the whole site. If the page returns a 404, your site has no robots.txt, which is the state 11.4% of the practices in this study are in.
- Open
yoursite.com/robots.txt. A 404 means there is no file, and no declared rules. - Find the
User-agent:lines. Each one starts a group of rules that applies to the crawler it names. - Under each group, look for
Disallow: /. That single slash closes the whole site to that crawler. ADisallow:with nothing after it blocks nothing at all. - Check the
User-agent: *group. Its rules apply to every crawler that does not have a group of its own, which is how sites block AI crawlers without ever naming one. - Decide crawler by crawler.
OAI-SearchBotandPerplexityBotgovern whether ChatGPT and Perplexity can show your practice.GPTBotandClaudeBotgovern whether your pages help train models.Google-Extendedcovers both Gemini training and grounding in Gemini apps, and Google says it does not affect Google Search.
If you would rather see the whole picture at once, the checker below reads your site and scores its AI readiness in seconds, crawler access included.
Prefer to question the engines by hand? See how to check what ChatGPT says about your practice, or the deeper AI visibility checker for dentists.
How we measured this
The sample frame is our own database of 7,632 US dental practice websites, one row per domain, collected from Google Maps business listings across 42 states. This run used the whole frame rather than a sample of it. Every figure on this page comes from that single pass on 22 July 2026, parsed by our own robots.txt engine, which ships with a 27-case self-test covering wildcard groups, specific-group overrides, empty files, malformed servers and the rest. It passed 27 of 27 when this page was written.
- A declaration, not enforcement. We read what the file says. A network or firewall layer can block a crawler the file allows, and we did not attempt to detect that here.
- 13 crawlers, not every crawler. Anthropic's Claude-SearchBot and Claude-User were not in this run, so the Anthropic figures cover training tokens only.
- Google Maps frame. Practices without a website, or not listed on Google Maps, cannot appear in the sample.
- Unreadable sites are excluded. 1,135 domains returned 401, 403, a server error, a DNS failure or a timeout, and were dropped from n rather than counted as open.
- A point in time. robots.txt files change. Every number here carries the 22 July 2026 scan date, and a rerun will produce different figures.
We publish a national figure only when we can read at least 2,000 sites, and a state or segment figure only at 300. The full framework behind our scans, scores and city reports is in how we measure AI visibility. If you think a number on this page is wrong, email contact@citevio.com with the domain and what you saw. We would rather be corrected than quoted wrongly.
Common questions
How many dental websites block AI crawlers?
In our July 2026 study of 6,497 US dental practice websites, 11.4% blocked at least one AI crawler at the root of the site. Narrowed to crawlers run by OpenAI, Anthropic or Perplexity, the figure is 7.6%, or about one site in thirteen. A further 11.4% had no robots.txt file at all, so their access rules were undeclared rather than closed.
Does blocking GPTBot remove my practice from ChatGPT search?
No. GPTBot is OpenAI's training crawler. The crawler that surfaces websites in ChatGPT's search results is OAI-SearchBot, and a rule aimed at one does not apply to the other. In our July 2026 study, 458 dental sites blocked GPTBot and 432 of them left OAI-SearchBot open. Only 26 sites blocked both.
Is having no robots.txt the same as allowing AI crawlers?
In practice it is close. RFC 9309, the robots.txt standard, says that if a server status code indicates the robots.txt file is unavailable to the crawler, then the crawler may access any resources on the server. So a missing file usually reads as open access. The difference is that nothing on the site records what the practice decided. In our July 2026 study, 739 of 6,497 dental sites, 11.4%, were in that state.
Should a dental practice block AI crawlers?
It depends on the crawler. Blocking GPTBot or ClaudeBot is a training opt-out and does not remove a practice from ChatGPT or Claude answers. Blocking OAI-SearchBot or PerplexityBot does remove it from ChatGPT search and Perplexity results. Google-Extended sits between the two: Google says it covers Gemini training and grounding in Gemini apps, and does not affect Google Search. Most practices want patients to find them, so the search-facing crawlers are the ones to leave open.