Open data · Free to download and reuse

Open data: how US dental practices look to AI

Every dataset behind our reports, free to download as CSV, with the scan dates, the denominators and the parts we will not publish. Built from our own scans, not bought.

This page is Citevio's data library. Three datasets are free to download as CSV: 527 dental practice websites read field by field in June 2026, 6,497 robots.txt files read in July 2026, and 499 homepages requested with seven crawler identities. Every file is aggregated, never row by row, so no practice is named.

Journalists, vendors and dental groups can use these files without asking. So can AI assistants, which is part of the point: a number is easier to cite when the file it came from is public and dated. What follows is one section per dataset, with what it measures, what it found, where to download it, and where it stops being reliable.

What is in this data library?

Three downloadable datasets and two worked city reports. The datasets cover how practice websites are built, what their robots.txt files allow, and what their servers actually do when an AI crawler asks for the homepage. The city reports apply the same method to one market at a time and are published as pages rather than files.

Each row below links to the method behind it. Every figure on this page comes from our own scans, and every scan window is stated, because a dated snapshot is honest and a live feed would be a claim we cannot support.

Dataset or reportWhat it measuresSizeScan windowFile
Dental practice web anatomyStructured data, platform, crawler rules, reviews, speed, Bing index527 practices6–26 June 2026CSV ×2
AI crawler access in robots.txtWhat robots.txt says about 13 AI and data crawlers6,497 sites22 July 2026CSV
Live crawler accessWhat the server does when an AI crawler asks for the homepage499 sites22 July 2026CSV
AI visibility report: CharlotteWhether ChatGPT and Perplexity name a practice for patient questions154 practices6–23 June 2026Report page
AI visibility report: ColumbusWhether ChatGPT and Perplexity name a practice for patient questions58 practices19–23 June 2026Report page

The two city reports have no CSV yet. Their underlying records are answer-by-answer readings tied to named practices, and we have not built an aggregate view of them that is worth publishing. When we do, it will appear here with the same rules as the rest.

The city reports and the June dataset were filtered separately, so their counts for the same city differ slightly. The Charlotte report was cleaned by hand and settled on 154 practices; the June dataset applies a written rule to every metro at once and keeps 153 in Charlotte. The one record between them is a listing whose only web address was a freelancer profile rather than a practice site, which the rule drops and the hand pass kept. The Austin report differs the other way: it settles on 111 practices, one more than this dataset's rule alone would keep, after we manually confirmed two pairs of listings on this dataset were the same practice under two different domains and folded each pair into one, and restored one nonprofit dental clinic the automated rule had dropped for lacking a dental word in its name or domain. Where a figure differs, the file on this page is the one we would defend for every metro we have not hand-checked, because it can be regenerated from the rule; Austin's report is the more carefully checked number for that one city.

What do dental practice websites actually look like?

Of 527 dental practice websites we read field by field in June 2026, 53.9% carried no LocalBusiness or Dentist schema at all, and only 3.4% had both LocalBusiness and FAQPage. Meanwhile 96.7% of those with a Google rating already sat at or above 4.3 stars. The reputation bar is cleared; the machine-readable layer is not.

This is the dataset we would point a reporter at first, because nothing in it is inferred. Each value is a field read straight off the site or its Google Business Profile, with no name extraction and no model judgment in the loop. The practices sit in seven metros, led by Charlotte at 153 and Austin at 112, plus 21 practices in nine smaller towns nearby. One row is one website: where two listings shared a site, they were merged, so a multi-site group counts once rather than once per location.

What we checkedFindingDenominator
Structured dataNo LocalBusiness or Dentist schema: 284 (53.9%) · LocalBusiness but no FAQPage: 225 (42.7%) · both: 18 (3.4%)527
Publishing platformWordPress 289 (54.8%) · custom or unidentified 187 (35.5%) · Wix 28 (5.3%) · Squarespace 22 (4.2%) · Shopify 1 (0.2%)527
robots.txtOpen to all six AI crawlers tested: 492 (93.4%) · no robots.txt file: 20 (3.8%) · blocks at least one: 15 (2.8%)527
Google ratingMedian 4.9 stars · at or above 4.3: 506 (96.7%) · at or above 4.5: 484 (92.5%)523
Review textMedian 250 reviews · median 77 words per review · median generic share 0.40523
Site speedMedian First Contentful Paint 3.01s · over 3s: 130 (50.8%) · over 4s: 75 (29.3%)256
Bing indexIndexed 256 (88.0%) · weak or thin 35 (12.0%)291

Source: Citevio scan of 527 US dental practices, 6–26 June 2026. Denominators differ by row because a practice with no completed speed reading is dropped from that row and no other. Method: how we measure AI visibility.

Every count in the first file reconciles: the metro rows for a given check add up to the row marked ALL, including the pooled small-town bucket. If they ever do not, the file is wrong and we would want to hear about it.

Why the star rating cannot be the differentiator

AI assistants are widely observed to favour practices above roughly a 4.3-star floor, and 96.7% of this sample already clears it. We have not measured that threshold ourselves and we do not present it as our finding. A signal that almost everyone passes cannot be what separates them. The variation sits further down: the median practice has 250 reviews averaging 77 words, and in the median practice about 40% of those reviews read as generic. Generic share is our own measure, defined in the methodology, not an industry metric.

How many dental sites block AI crawlers in robots.txt?

Fewer than the headline suggests. Of 6,497 US dental practice sites whose robots.txt we could read on 22 July 2026, 11.4% block at least one of the 13 AI and data crawlers we tested at the root. Counting only crawlers run by OpenAI, Anthropic and Perplexity, it falls to 7.6%. Another 11.4% publish no robots.txt at all.

The two numbers are not interchangeable, and the gap is mostly one crawler. Bytespider, run by ByteDance, is blocked by 11.1% of the sample on its own and is not used by any assistant a patient would ask. OAI-SearchBot, which OpenAI documents as the crawler that surfaces sites in ChatGPT's search results, is blocked by 0.4%.

The eight most-blocked of the 13 crawlers tested. All 13, including anthropic-ai, ChatGPT-User, Perplexity-User, Applebot-Extended and Bingbot, are in the CSV.
CrawlerWho runs itBlocked at the root
BytespiderByteDance722 (11.1%)
ClaudeBotAnthropic484 (7.4%)
meta-externalagentMeta483 (7.4%)
CCBotCommon Crawl471 (7.2%)
GPTBotOpenAI (model training)458 (7.0%)
PerplexityBotPerplexity272 (4.2%)
Google-ExtendedGoogle (Gemini training)223 (3.4%)
OAI-SearchBotOpenAI (ChatGPT search)26 (0.4%)

Source: Citevio robots.txt study, 22 July 2026, n = 6,497 readable domains of 7,632 fetched. Crawler roles as documented by OpenAI, Anthropic and Perplexity. Full method: are dental websites blocking AI crawlers?

The file carries state cuts for the four states where we could read at least 300 sites. California blocks at the highest rate of the four, 15.0% of 1,751 sites on the 13-crawler measure and 11.6% on the assistant-only measure, against 7.0% and 4.8% of 1,151 sites in Texas. Below 300 readable sites we publish no state figure, because a rate on a thin base moves with a handful of websites.

It also carries cuts by the business category Google shows on the listing, again only where at least 300 sites were read: Dentist (3,929 sites, 12.5% blocking), Cosmetic dentist (756, 11.0%), Orthodontist (415, 4.8%) and Dental clinic (382, 11.0%). Read that Cosmetic dentist row as what a listing is labelled, not as a verified clinical specialism, and do not read it as a cosmetic-dentistry sample.

A missing robots.txt is not a block. RFC 9309, the standard that defines the file, states that "if a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server." For the 11.4% with no file, what is lost is control rather than access.

Source: RFC 9309, Robots Exclusion Protocol, IETF, 2022.

Do servers honour what robots.txt promises?

Often not. Of 499 dental sites whose robots.txt welcomes AI crawlers, 14.4% were refused by their own server anyway: 69 sites returned an outright block status such as 403 or 429, and three returned a bot challenge page. In 12.6% of the sample, Googlebot was served normally in the same run.

That last figure is what makes this a finding rather than a bot wall. Only 6 sites of 499 refused Googlebot at all. The rule the file publishes and the rule the server enforces are two different rules, and the gap falls on the crawlers that feed AI answers.

Crawler identityRefused byShare of 499
ClaudeBot46 sites9.2%
GPTBot40 sites8.0%
OAI-SearchBot17 sites3.4%
PerplexityBot10 sites2.0%
ChatGPT-User7 sites1.4%
Bingbot6 sites1.2%
Googlebot (control)6 sites1.2%

Source: Citevio live access study, 22 July 2026. 550 domains drawn at random from a frame of 5,019 whose robots.txt does not block AI crawlers at the root; 499 remained after dropping sites where a plain browser request also failed. Full method: open to AI on paper, blocked in practice.

A further 41 sites, 8.2% of the sample, served every crawler but sent a different page body to at least one of them. We publish that count as a flag rather than a finding, because a different body can mean cloaking or it can mean an ad slot that rendered differently on one request.

What did we remove before publishing?

Three kinds of detail come out before a file is written: the practice domain column, every row-level record, and any cell too thin to publish safely, whether that is a town with fewer than 20 scanned practices or a state cut built on fewer than 300 sites. A fourth removal is different in kind: 84 scanned records were dropped from the June set because they were not dental practices at all, and repeat scans of the same website were merged into one row. What remains is counts, and no row in any published file describes a single business.

Our raw scan files name practices and pair them with technical faults. Publishing that list would make a public register of clinics that are, for example, invisible to Bing, and no practice agreed to be on it. The value of the data is in the pattern, and the pattern survives aggregation intact.

  • Domains, names and URLs are dropped before a file is written, not hidden in it.
  • Small towns are pooled. Nine places with fewer than 20 scanned practices appear as one bucket of 21, so a town with a single practice cannot be read as that practice.
  • Thin cells are not published. State and category breakdowns need at least 300 sites; metro breakdowns need at least 20 practices.
  • Repeat scans of the same website were merged. The raw scan folder held 21 websites read more than once, sometimes under two spellings of the practice name. Each is now one row: the most recent scan, except where that scan failed on a given check, in which case the most recent working reading of that check is used. A failed check is not a result.
  • The same practice under two different domains is not always caught. Merging above works by matching domains, so it cannot catch one practice running two separate websites under two different names. We found and manually confirmed two such pairs in the Austin set (identical rating, review count and business identity, verified by opening both sites), which is why the Austin city report on this site (111 practices) differs slightly from what a domain-only match on this dataset would produce (112). We have not run that same manual check against every metro; a handful of other exact rating-and-review-count matches turned up in a spot check and have not been confirmed one way or the other, so the published per-metro counts may be very slightly high for this reason.
  • 84 records were dropped from the June set before publication because they were not dental practices. The list we scanned had collected furniture stores, salons and plastic surgeons along the way; four dental laboratories, which manufacture for practices rather than treat patients, and one listing whose only web address was a freelancer profile were dropped for the same reason. Anything without a dental signal in its business name, its domain or its listing category was removed, which also drops a few real dental clinics whose names give nothing away. That leaves 527 practice websites from 611 unique sites scanned, and it is the reason a figure here can differ slightly from an earlier Citevio note built on the unfiltered set.
  • One record was out of window. A single scan ran on 6 July, ten days after the stated window closed, in a metro nowhere near the other seven. Rather than widen the published dates for one row, we dropped it.

What is deliberately missing from these files?

Four measurements sit in our files and not on this page. The Foursquare listing check is wrong at the matcher. The study of which practices AI assistants name is not yet reliable. A clinical cosmetic-dentistry breakdown does not exist, because the field that would mark it is absent from every scan file we hold. And the Bing index reading is published for the sample as a whole but not by metro.

Saying so is cheaper than being corrected later. Each of these will be published when the instrument is fixed, with the date of the rerun on it.

Held backWhyWhat would release it
Foursquare listing checkThe matcher returned hotels and gyms for dental queries, so the counts describe the wrong businessesA rewritten matcher and a rerun
Which practices AI namesThe step that pulls names out of an answer still yields phrases like "Winner Best Dentist", and the repeat runs sat inside one session rather than across daysNames stored per engine and query, repeats spread over days
Cosmetic-only breakdownThe field that marks a practice as clinically cosmetic is in the scanner but in none of the scan files we hold. The July file's "Cosmetic dentist" rows are Google's listing label, which is not the same thingThe next scan round
Bing index by metro236 of 527 index checks failed on our side, which leaves the per-metro base too thin to publishA rerun without the API rate limits

The Bing line is worth reading twice, because it is the trap in this kind of work. A failed check looks exactly like a bad result unless you read what the failure says. In the June set, 236 index checks came back as failures from our own data provider: 220 rate-limit responses and 16 read timeouts. Counting those as "weak in Bing" would have produced a much larger and completely false number, so the published figures use only readings that came back.

How do I cite and reuse this data?

Freely, including in commercial work, if you credit Citevio and link to this page. Say which scan date you used, because these are dated snapshots and not a live feed. If you want a cut for your own readers, a single state or a single crawler, email contact@citevio.com and we will run it.

Ready-made citation lines, one per dataset:

Citevio, dental practice web anatomy, June 2026 (n = 527). citevio.com/data
Citevio, AI crawler access in dental robots.txt files, 22 July 2026 (n = 6,497). citevio.com/data
Citevio, live crawler access on dental homepages, 22 July 2026 (n = 499). citevio.com/data

Each CSV is tidy: one row per metric, segment and breakdown, with the count, the denominator and the percentage in their own columns, plus a plain-language definition of what that denominator holds. That last column matters more than it sounds, because the denominators genuinely differ between rows and quoting a percentage without it is how our own summaries went wrong before.

What are the limits of this data?

Six. The June set covers seven metros and is not a national sample. The speed reading covers 256 of 527 practices. One row is one website, not one location. The crawler studies describe one day in July. Our two crawler instruments give two different blocking rates. And all of it measures what a site or a server does, not why an assistant chose one practice over another.

  • Sample, not census. The June practice set is 527 practices in seven metros plus nine nearby towns. It supports "in this sample" and does not support "nationally". The July robots.txt study is wider, covering 42 states, but it is not evenly spread: California and Texas alone supply about 45% of the domains we could read, which is why we publish state figures separately instead of leaning on the national average.
  • Speed has a partial base. First Contentful Paint completed for 256 of 527 practices. In the rest the field was still marked as pending rather than failed, so we cannot say whether the missing sites are faster or slower. The direction of that bias is unknown, and we will not guess it.
  • One website, one row. Practices sharing a website are merged, so a group running several locations off one site appears once. That is the right unit for a study of websites and the wrong one for counting locations, and the review figures for a merged row describe the listing we read, not the group.
  • One day, one instrument. Both crawler studies were run on 22 July 2026. Cloudflare and similar services change bot rules often, and a rerun will not reproduce the file exactly.
  • Two instruments, two rates. The June set checks six crawler names and finds 2.8% blocking at least one; the July study checks 13 and finds 11.4%, or 7.6% for assistant crawlers only. On the 456 domains present in both, the readings agree on 432, or 94.7%. Different samples and crawler lists explain the rest, and the July study is the better instrument for the blocking question.
  • Description, not cause. Nothing here tests whether a slow site or a missing schema causes an assistant to skip a practice. We have not measured that, so we do not claim it.

For how the scans themselves are run, which engines we test and what counts as visible, see how we measure AI visibility. For what Citevio is and who is behind it, see what is Citevio.

See where your own practice stands

These files describe a sample. The checker below reads your site and scores its AI readiness against the same categories, in seconds, with no email and no call.

Prefer the deeper version? See the AI visibility checker for dentists.

Frequently asked questions

Can I republish Citevio's data?

Yes, including in commercial work. Credit Citevio and link to citevio.com/data, and say which scan date you used, because these files are dated snapshots rather than a live feed. A citation line such as "Citevio, dental practice web anatomy, June 2026 (n = 527)" is enough. If you want a cut of the data for your own readers, such as a single state or a single crawler, email contact@citevio.com and we will run it.

How are these datasets anonymized?

The published files hold counts, not records. Practice domains, names and URLs are removed before a file is written, so no row describes one business. Towns with fewer than 20 scanned practices are grouped into a single bucket, and breakdowns by state or business category are published only where at least 300 sites were read. Citevio holds the row-level data and does not publish it, because a public list of named practices with technical faults is a list nobody consented to be on.

Why do two Citevio datasets give different AI crawler blocking rates?

Because they are different samples read with different instruments a month apart. The June diagnostic set covers 527 practices in seven metros and checks six crawler names, and 2.8% of those sites block at least one of them. The July study covers 6,497 sites nationwide and checks 13 crawler names, and 11.4% block at least one, or 7.6% counting only crawlers run by OpenAI, Anthropic and Perplexity. On the 456 domains that appear in both, the two readings agree on 432, or 94.7%.

What does Citevio measure but not publish?

Four things. The Foursquare listing check is withheld because its matcher returned hotels and gyms for dental queries, so its numbers are wrong. The study of which practices AI assistants name is withheld because the step that pulls names out of an answer still produces items like "Winner Best Dentist", which is a phrase and not a practice. A clinical cosmetic-dentistry breakdown does not exist, because the field that would mark a practice as cosmetic is absent from every scan file we have; the July robots.txt file does carry a "Cosmetic dentist" cut, but that is the category Google's business listing shows, not a verified clinical specialism, and the two should not be read as the same thing. And the Bing index reading is published for the whole sample but not by metro, because 236 of the 527 index checks failed on our side.

How often is this page updated?

Each dataset carries its own scan window, and that window is the honest read of how fresh it is. When a scan is repeated, the CSV is replaced and the date on this page changes with it. The June 2026 practice set and the July 2026 crawler studies were run once each, so a figure taken from them describes those days and not today.