- 0:00 The number: 11.4% and what it actually means
- 0:30 Two different studies · why we're keeping them apart
- 1:11 Methodology: 7,632 domains requested, 6,497 readable
- 1:42 What robots.txt actually is
- 2:19 Finding 1 · 11.4% block at least one of 13 crawlers
- 3:09 Finding 2 · narrowed to OpenAI, Anthropic, Perplexity: 7.6%
- 4:41 Most sites block nothing · so why still invisible?
- 5:01 A separate 527-practice scan: still 96.2% unmentioned
- 5:46 What we're NOT claiming (twice · access isn't causation)
- 8:05 Recap and where the open data lives
What is the finding, in one paragraph?
In July 2026, robots.txt was read on 6,497 of 7,632 requested US dental practice domains. 11.4% blocked at least one of 13 tested AI and data crawlers; 7.6% when narrowed to only OpenAI, Anthropic and Perplexity's own crawlers. A separate June 2026 scan of 527 dental practices, most already open to AI crawlers, found 96.2% got zero mentions from ChatGPT or Perplexity. Access is a precondition, not a cause.
What was the methodology?
Robots.txt was requested from 7,632 dental practice domains nationally in July 2026. 6,497 came back readable; the remainder hit a server error, a timeout, or returned nothing, and were dropped rather than counted as blocking. A domain that couldn't be checked isn't a domain that's blocking anyone.
For a stable citation to the two datasets used here, the archive identifier is DOI 10.5281/zenodo.22016876.
What is robots.txt?
Robots.txt is a note published at the root of a website listing which automated visitors are welcome. OpenAI, Anthropic, Perplexity and Google each publish which of their crawlers look for that note. A rule-following crawler reads the note first; if it sees a disallow with its name on it, it walks away without requesting anything else.
Crawler documentation: OpenAI, Anthropic, Perplexity, Google.
What are the two numbers, and why aren't they interchangeable?
11.4% of readable sites block at least one of 13 tested AI and data crawlers, a list spanning training crawlers, answer-generating crawlers, and some crawlers no assistant a patient uses is likely to run. Narrowed to only the crawlers OpenAI, Anthropic and Perplexity themselves operate, the rate drops to 7.6%. The two figures measure different things and are never used interchangeably in this research.
| Measure | Sample | Result |
|---|---|---|
| Blocks ≥1 of 13 tested crawlers | 6,497 readable domains | 11.4% |
| Blocks an OpenAI/Anthropic/Perplexity crawler | 6,497 readable domains | 7.6% |
| Zero mentions from ChatGPT or Perplexity | 527 practices (separate June 2026 scan) | 96.2% |
Why doesn't an open door mean a mention?
In a separate scan of 527 dental practices across seven US cities in June 2026, robots.txt was already open to AI crawlers for nearly every practice, and 96.2% (507 of 527) still received zero mentions from ChatGPT or Perplexity. An open robots.txt file does not cause an AI assistant to mention a practice; being reachable is a precondition, the practice still has to look like the right answer to whatever question was asked.
What doesn't this research measure?
This study reports what a site's robots.txt file declares. It does not test whether a server actually honors that file when a real crawler requests a page. That is a different question, examined separately in another study on this channel, because the two don't always agree.
Scope note: this scan tested ClaudeBot and the older anthropic-ai token, not Claude-SearchBot or Claude-User, so the Anthropic figures above measure training opt-outs only, not search access. Source: dental-websites-blocking-ai-crawlers, "How we measured this".
Citevio's dental market research is published openly under CC-BY-4.0 with a DOI, so the numbers can be checked and reused; client data is never published.
What are we not claiming?
We are not saying blocking is common: for most sites, it isn't. We are not saying an open file gets a practice recommended: on its own, it doesn't. And we are not saying robots.txt is the reason most dental practices are invisible to AI, because nothing in either dataset supports that claim.
What the data does point to instead is covered in the shared trait among the practices AI does skip.
Frequently asked questions
Do most dental practices block AI crawlers in robots.txt?
No. In a July 2026 study of 6,497 readable dental practice robots.txt files, only 11.4% blocked at least one of 13 tested AI and data crawlers, and only 7.6% blocked a crawler that OpenAI, Anthropic or Perplexity themselves operate. The large majority block nothing.
Does an open robots.txt mean ChatGPT or Perplexity will mention a dental practice?
No. A separate scan of 527 dental practices found that even with robots.txt open to AI crawlers for nearly all of them, 96.2% still received zero mentions from ChatGPT or Perplexity. Crawler access is a precondition, not a guarantee of a mention.
What's the difference between the 11.4% and 7.6% figures?
11.4% is the share of readable dental sites blocking at least one of 13 tested AI and data crawlers. 7.6% narrows that to only the crawlers OpenAI, Anthropic and Perplexity themselves operate. The two are not interchangeable and measure different scopes of blocking.
How many dental practice domains were requested for this robots.txt study, and how many were actually readable?
7,632 dental practice domains were requested nationally in July 2026, and 6,497 came back readable. The remainder hit a server error, a timeout, or returned nothing.
What happened to the domains whose robots.txt file couldn't be read?
They were dropped from the study rather than counted as blocking. A domain that couldn't be checked isn't a domain that's blocking anyone, so counting it as a block would have inflated the finding.
What is robots.txt, in plain terms?
Robots.txt is a note published at the root of a website listing which automated visitors are welcome. OpenAI, Anthropic, Perplexity and Google each publish which of their crawlers look for that note, and a rule-following crawler reads it first — if it sees a disallow with its name on it, it walks away without requesting anything else.
Why does this study test 13 different crawlers instead of just two or three?
Because the 13 crawlers tested don't do the same job. The list spans crawlers that gather training data, crawlers that power a company's own search product, and a few that don't belong to any assistant a patient is likely to be using at all — treating them as one number would hide those differences.
If a site blocks one of the 13 tested crawlers, does that mean it's blocking the crawler that actually powers a patient's AI answer?
Not necessarily. A site can block a crawler that only collects training data while leaving every assistant's actual answer-generating crawler wide open, or the other way around — blocking one of the 13 doesn't tell you which situation a specific site is in.
What's the difference between OpenAI's and Anthropic's training crawlers and their other crawlers?
OpenAI and Anthropic each run a training crawler and a separate crawler that's closer to answering a live question. Perplexity's crawler is the one that surfaces and links pages in its results, not a training tool at all.
Does the 7.6% figure tell me which specific crawler a blocking practice is blocking?
No. The 7.6% figure — sites blocking a crawler that OpenAI, Anthropic or Perplexity themselves operate — doesn't say which crawler any single site is blocking, so a practice could be blocking only a training crawler and never affect whether it gets cited, or block the one crawler that actually matters and quietly opt out of being cited at all.
How many of the 527 practices in the separate June 2026 scan already had robots.txt open to AI crawlers?
Nearly every one of the 527 dental practices in that separate June 2026 scan already had robots.txt open to AI crawlers.
How many of those open-robots.txt practices still got zero mentions from ChatGPT or Perplexity?
507 of 527 (96.2%) of those practices still received zero mentions from ChatGPT or Perplexity, despite having an open door.
Did this study check whether a dental practice's server actually honors its own robots.txt file?
No. This study reports what a site's robots.txt file declares, not whether the server actually honors that file when a real crawler requests a page — that's a different question, examined separately in another study, because the two don't always agree.
Does this study's Anthropic figure include Claude's search crawlers, or only its training crawler?
Only its training crawler. This scan tested ClaudeBot and the older anthropic-ai token, not Claude-SearchBot or Claude-User, so the Anthropic figures measure training opt-outs only, not search access.
Is the narration in this video real, or is it synthesized?
The narration is synthesized. The data, methodology and limitations discussed in it are Citevio's own and are linked from the page.
Where can I check the source data behind this video myself?
Every number in this video, the full methodology, and both underlying datasets are published at citevio.com/data.
Does this video claim robots.txt is why most dental practices are invisible to AI?
No. The video is explicit that it isn't saying robots.txt is the reason most dental practices are invisible to AI, because nothing in either dataset it uses supports that claim.
If my robots.txt already looks fine, is there still something worth checking?
Yes — checking is still worth five minutes. Open your own domain, add /robots.txt, and read what's there for a line that blocks everything or names a crawler you don't recognize; an open file is a starting point, not proof that your practice is actually being found.
Where does this research say the real explanation for AI invisibility lies, if not robots.txt?
That's covered separately, on the page examining why AI search ignores most dental practices.
What topics does this video walk through, in order?
From the headline number and what it means, through the two-study methodology, the 11.4% and 7.6% findings, the comparison against the 527-practice mention scan, what the study is explicitly not claiming, and a recap of the source data.
Full transcript
Here's a question most dental practice owners assume they already know the answer to. If an AI assistant can't see your website, are you the one blocking it? We went and checked, at national scale. 11.4% of dental websites block at least one AI or data crawler in their robots.txt file. That leaves the large majority not blocking anything — which is exactly where this video gets interesting.
One thing before we go further. This is a separate study from the 527-practice scan we've published elsewhere on this channel. Different month, different sample, different question. That earlier scan asked whether ChatGPT or Perplexity would actually name a practice. This one asks something narrower and more mechanical: nationally, what does a dental website's robots.txt file say about AI crawlers? Keep the two apart — we'll bring one number back in for comparison later, and we'll say so clearly when we do.
In July 2026, we pulled robots.txt from a national list of dental practice domains — 7,632 of them. 6,497 came back readable. The remainder hit a server error, a timeout, or simply went quiet, and we dropped those instead of guessing at what they meant. A domain we couldn't check isn't a domain that's blocking anyone, and treating the two the same would inflate this finding.
Think of robots.txt as a note taped to a website's front door, listing which visitors are welcome. It lives at the root of the site, and a handful of companies — OpenAI, Anthropic, Perplexity, Google — tell the public exactly which of their automated readers are looking for that note. A crawler that plays by the rules reads the note first. See a no with its name on it, and it walks away without asking for anything else. Small file, real consequences — which is the entire premise behind this video.
Here's the headline number, in full. Run the check across every readable site in that pool of 6,497, and 11.4% of them turn away at least one of the thirteen AI and data crawlers we tested. Thirteen, not two or three — the list includes crawlers that gather training data, crawlers that power a company's own search product, and a few that don't belong to any assistant a patient is likely to be using at all.
Blocking one of those thirteen doesn't mean the same thing every time. A site can block a crawler that only collects training data while leaving every assistant's actual answer-generating crawler wide open — or the other way around. That difference is the entire reason we didn't stop at one number.
So we narrowed it. Drop everything except the crawlers that OpenAI, Anthropic and Perplexity themselves operate — the three companies behind ChatGPT, Claude and Perplexity. Narrowed that way, the block rate drops to 7.6%. Roughly one dental website in thirteen.
Even inside that narrower group, not every crawler does the same job. OpenAI and Anthropic each run a training crawler and a separate one that's closer to answering a live question. Perplexity's crawler is the one that surfaces and links pages in its results, not a training tool at all. A practice can block a training crawler out of caution and never affect whether it gets cited — or block the wrong one and quietly opt out of being cited at all. The 7.6% doesn't tell you which situation any single site is in.
Two numbers, then: 11.4% and 7.6%. They are not interchangeable, and we're going to say the qualifier every time we use them. 11.4% blocks at least one of thirteen tested crawlers. 7.6% blocks specifically a crawler that OpenAI, Anthropic or Perplexity runs. If you see either number repeated without that phrase attached to it, the qualifier got lost — not the finding.
Turn the number around and the more useful fact is the one we didn't put in the title: the large majority of the dental websites we could read aren't blocking any AI crawler at all. Whatever is keeping most dental practices out of AI answers, for most of them, it is not a locked door.
So does an open door mean an AI assistant actually mentions you? We can check that against a different, smaller sample. In a separate scan of 527 dental practices across seven US cities this past June, robots.txt was already open to AI crawlers for nearly every one of them. And in that same 527-practice sample, 96.2% — 507 of 527 practices — still got zero mentions from ChatGPT or Perplexity. Different sample, different month, and the shape repeats anyway: an open door did not translate into a mention.
Say that plainly, because it's easy to slide past. An open robots.txt file does not cause an AI assistant to mention a practice. Being reachable is a precondition, not a guarantee — the practice still has to look like the right answer to whatever question got asked. The 96.2% figure from that other scan is exactly why we're not claiming otherwise.
Run the logic the other direction and the same caution applies. A blocked crawler doesn't explain most of this gap either. Only 11.4% of the sites in this study block anything, and narrowed to the crawlers that actually power an assistant's answers, it's 7.6%. That's not a large enough share to account for how rarely dental practices show up in AI answers overall. Where blocking happens, it's a real problem — it just isn't the main story.
One more limit, and we'd rather flag it than let you assume otherwise. We read the file a site publishes. We didn't test whether a server actually honors that file when a real crawler shows up asking for a page — that's a different question, and we looked at it separately in another study on this channel, because the two don't always agree. This page only tells you what a practice declared, not what happens at the door when someone actually knocks.
If you run a dental practice site, the file is worth five minutes. Open your own domain, add slash robots.txt, and read what's there. Look for a line that blocks everything, or that names a crawler you don't recognize. If nothing there says no to OpenAI, Anthropic or Perplexity's crawlers, this particular door is already open — and now you know that being open is not the same as being found.
To be clear about what this video is not saying. We're not saying blocking is common — for most sites, it isn't. We're not saying an open file gets a practice recommended — on its own, it doesn't. And we're not saying robots.txt is the reason most dental practices are invisible to AI, because nothing in either dataset supports that claim.
So, to recap. We read robots.txt on 6,497 dental websites out of 7,632 requested. 11.4% block at least one of thirteen tested AI and data crawlers. Narrowed to just the crawlers OpenAI, Anthropic and Perplexity run themselves, that's 7.6%. And in a separate, earlier scan of 527 dental practices, 96.2% still got zero mentions from ChatGPT or Perplexity — despite most of them already having the door open.
Every number in this video, the full methodology, and both datasets are open at citevio.com/data. We'd rather be corrected than quoted wrong — if a figure here looks off, the contact is right there on the page. This channel publishes measured, dated research on how AI search engines find, or don't find, dental practices — not guaranteed rankings. Thanks for watching.
Narration in this video is synthesized; the data, methodology and limitations are our own and are linked below.