Methodology · Measuring AI visibility

How AI visibility is measured

AI visibility is whether an AI assistant names a business when someone asks it for a recommendation. Measuring it means fixing five things in advance: the list of patient-style questions you test, the AI engines you ask, the date you ask them, what counts as "visible" in the answer, and how many times you repeat each question before trusting the result.

We test the same patient-style questions across ChatGPT, Perplexity, Gemini and Google AI Overviews, run each one several times because answers shift, and score five things: crawler access, structured data, Bing indexing, reputation signals, and what the engines actually say. The framework on this page is public; the exact weighting behind the score is ours.

This page is the method that our city reports point back to, so you can check our work. It explains what we count, how we count it, and, just as plainly, where the method has limits. Nothing here is a guarantee about your practice. It is a description of how we look.

What do we measure?

We measure one thing: how often an AI assistant recommends a practice to a patient, and how ready that practice is to be recommended. We test four engines patients actually use to pick a dentist, with the phrases patients actually type, then record whether each engine names the practice, mentions it, or leaves it out.

The reason this is worth measuring is how few names each assistant gives. In its 2026 Local Visibility Index of roughly 350,000 business locations, SOCi found that only 1.2% of locations get recommended on ChatGPT and 7.4% on Perplexity, against 11% on Gemini and 35.9% in Google's local 3-pack. When an answer holds a short list of names instead of ten links, being measured accurately matters more, not less.

Source: SOCi Local Visibility Index, 2026, via Search Engine Land.

We test the queries a real patient would ask, swapping in the city and neighborhood. Typical examples are best Invisalign dentist in [city], cosmetic dentist [city], and [city] dentist open Saturday. Practice-name searches are treated separately, because they only prove a clinic exists, not that it gets recommended.

A practice's set is split into five parts: a fixed core carried over from our published scans, questions for the treatments the practice actually offers, decision-stage questions, five questions that name the practice directly to catch what engines get factually wrong about it, and — on the top plan — competitor comparisons. The named-practice questions are tracked separately and are not what the citation guarantee is judged on: being named in an answer to your own name shows you exist, not that you get found.

This is also how we count a citation. A mention only counts if it came back for a commercial-intent question — one where a patient is choosing a treatment, a place, a price or a timeline. An engine quoting your practice inside a general-knowledge answer, or naming you because someone typed your name, is recorded but not counted: it tells you the engine knows you exist, not that it puts you forward when a patient is deciding.

Citevio's published method is the Citation-to-Chair Protocol, six named layers — Discover, Read, Match, Trust, Answer, Chair — each targeting one break in the path between an AI assistant naming a practice and a booked case. The protocol runs on continuous measurement: between 3,600 and 7,200 AI answers are logged every month depending on plan, and every layer decision follows what those measurements show. Read the Citation-to-Chair Protocol

What counts as "visible"?

We sort each result into three states. Recommended means the engine puts the practice forward as an answer to the patient's question. Mentioned means the practice appears somewhere, such as a list or a cited page, but is not put forward as the recommendation. Not visible means it does not appear at all. Only the first state wins a patient.

The distinction matters because it is easy to feel visible and still lose the patient. A clinic can be mentioned inside a directory the engine cites and never be the name the assistant actually says. So we count a practice as visible for a query only when the engine names it in its own answer, not when it merely sits on a page somewhere in the sources.

StateWhat it meansDoes it win the patient?
RecommendedThe engine names the practice as an answer to the patient's questionYes
MentionedThe practice shows up in a list or a cited source, but is not the recommendationRarely on its own
Not visibleThe practice does not appear for the query at allNo

How do we run scans?

For each practice we run a fixed set of patient-style queries across each engine, and we run every query several times over a short window rather than once. We record who gets named each time. Because a single answer is a snapshot, the number that matters is how consistently a practice appears, not whether it showed up on one lucky run.

Running once would be misleading, because AI answers are rebuilt each time from a shifting set of sources. According to a 2025 Profound analysis of about 80,000 prompts per engine, roughly 40–60% of the domains an engine cites for a given question are different one month later. That is why we treat visibility like weather readings: several samples, then the pattern, never a single day.

Source: Profound domain-drift analysis, 2025 (~80,000 prompts per engine, June–July). tryprofound.com

Where does our data come from?

Every figure in our reports comes from our own scans, not bought datasets. In June 2026 we scanned dental practices across seven US metros: Charlotte, Columbus, Austin, Raleigh, Nashville, Tampa and Salt Lake City. Each report shows the exact dates its scan was run, so you can judge how fresh the picture is before you rely on it.

Building the data ourselves is the point. It is information no competitor can copy, and it keeps us honest, because we publish the scan window instead of implying the data is live. You can see a full worked example in our AI visibility report for Charlotte dentists, which applies this exact method to one market. When a city is re-scanned, that report is updated and dated again.

The publication rule follows the same auditability principle. Citevio's dental market research is published openly under CC-BY-4.0 with a DOI, so the numbers can be checked and reused; client data is never published.

How do we measure change over time?

With a panel: the same websites read more than once, rather than a fresh sample each time. Our robots.txt time series follows 2,895 US dental practice websites read three times — in an Internet Archive copy dated before August 2023, in one dated after it, and live on 22 July 2026. Because every site appears in all three readings, a change in the rate is the same practices changing their rules, not a different sample answering a different question.

August 2023 is the dividing line because OpenAI published GPTBot and its opt-out instructions that month. We fixed that cut before looking at the results, so it is a prior taken from a public event rather than a line drawn around a finding.

A site only enters the panel if the archive can answer for it in both windows. We attempted the full frame of 5,823 domains and report exactly what happened to the rest, because the sites an archive cannot answer for are part of the method, not a rounding error.

Attrition on the robots.txt time series. Every domain in the frame was attempted; none was sampled out.
OutcomeDomains
In the panel (both windows usable)2,895
No archive coverage at all1,145
Both requests returned the same snapshot918
Only a pre-cut snapshot exists ("after" too old)761
Only a post-cut snapshot exists ("before" too recent)104

Archived files are re-parsed from their raw text with two independent parsers: Python's standard urllib.robotparser and a purpose-written RFC 9309 implementation. The two disagree on 0.47% of parsed files (25 of 5,339), and on the strict parser the series reads 1.3% to 4.3% instead of 1.6% to 4.4% — conservative about the size of the rise, not the direction. We publish the more conservative pair. A finding that changed depending on which parser we chose would not be a finding.

We also count crawlers the same way in every window. The live study tests thirteen crawler names and the archive pipeline eighteen; the time series counts only the twelve on both lists, so no part of the change can come from a longer list being applied to the newer data.

Then we try to break it. If practices had merely started publishing a robots.txt at all, the rise would be an artefact — restricting to the 2,504 sites with a file in both windows gives the same shape. If one crawler on a copied blocklist explained it, dropping Bytespider would flatten it — on assistant crawlers alone it still rises. And if these were sites simply going dark, they would be blocking Googlebot too: of the 127 blocking an AI crawler in the later window, none did. And if the jump were an artefact of switching instruments — archived text for the first two windows, a live scanner for the third — the two would disagree on live data: re-fetching 149 domains and reading them with both the archive parser and the live scanner, the two agreed on 148 of 149; the one exception is a site whose robots.txt now returns a different response than it did in July, a site-side change in the four weeks between checks, not a parser disagreement.

Full figures, confidence intervals and the downloadable CSV: Citevio's data library.

For a permanent identifier rather than a mutable page URL, use DOI 10.5281/zenodo.22016876.

What does each AI engine read?

The four engines do not share one source of truth, so we check each on its own terms. ChatGPT reads Bing's web index. Perplexity reads the live web and shows its citations, favoring fresh pages. Gemini and Google AI Overviews lean on Google's index and your Google Business Profile. A practice can be strong in one and absent from another, which is why a single-engine check is never enough.

The table maps each engine to the signals it builds local answers from, and to the part of our score that tests for it. It is also why our checker queries all four rather than assuming ChatGPT speaks for every assistant a patient might use.

EngineMain signals it builds local answers fromWhat our score checks for it
ChatGPTBing's web index, plus consistent business details and mentions across platformsBing indexing and reputation signals
PerplexityThe live web with visible citations; a preference for fresh, well-structured pagesCrawler access and structured data
Gemini / Google AIGoogle's index plus your Google Business ProfileReputation signals and the AI answer test

Sources: ChatGPT's reliance on Bing per Damian Rollison, Search Engine Land, 2025; Gemini's use of Google's index and Business Profile per SOCi's 2026 Local Visibility Index; Perplexity's visible-citation design is observable in the product.

What does our score mean?

Our score runs 0 to 100 and rolls up five categories: crawler access (can AI bots reach the site), structured data (can engines read the practice correctly), Bing indexing (is the site actually in Bing), reputation signals (reviews, ratings and consistent details), and the AI answer test (what the engines say when asked). We publish these categories. We do not publish the weighting between them.

The crawler-access category is the one we have measured at national scale. Two studies sit behind it: are dental websites blocking AI crawlers?, which reads robots.txt across 6,497 US dental practice websites, and open to AI on paper, blocked in practice, which requests the homepage with seven crawler identities to see whether the server honours what the file promises.

We have not measured whether reach, reputation or any technical change causes better AI visibility outcomes. Correlation and citation-distribution studies can describe associations and source patterns; they do not establish which change caused an engine to name a practice.

Sources: Ahrefs study of 75,000 brands, 2025 (ahrefs.com); Muck Rack "What is AI Reading", May 2026 (muckrack.com).

One deliberate limit: structured data sits in the score as reading infrastructure, not a citation trigger. Schema helps an engine read your name, address, hours and services correctly. It does not, on its own, make an engine cite you. We score it for what it does, and no more.

A note on reputation signals. The locations ChatGPT recommended in SOCi's 2026 index averaged 4.3 stars, and they weigh how consistent and recent your reviews are. That is why two clinics with the same technical setup can score differently: the one patients rate and describe more often reads as the safer recommendation.

What are the limitations?

This method has real limits, and hiding them would make the data less useful. AI answers change from run to run, so even several samples are a recent picture, not a permanent one. Some scans check Gemini manually rather than through the same automated pipeline. And a scan measures what an engine says, not why, so we report visibility, not a guaranteed cause behind it.

  • Answers move. With 40 to 60% of cited domains changing monthly (Profound, 2025), a practice that is visible today may not be next month, and the reverse.
  • Gemini is sometimes manual. Where a scan did not include Gemini in its automated column, we checked it by hand and say so in the report rather than inventing a number.
  • It is a point in time. Each report carries its scan dates. We do not present month-old data as if it were live.
  • We measure, we don't promise. Higher scores line up with getting recommended, but no honest method can guarantee a spot in an answer no one controls.
  • The time series is a surviving panel. A site enters it only if the Internet Archive covered it in both 2022 and 2025 and it was still reachable in 2026, and archive coverage favours better-linked sites. The panel reads 12.3% blocking where the full 6,497-site study reads 11.4%. Cite the direction and size of the change; treat the level as approximate, and read nothing in it about practices that opened after 2022.
  • A rule is not a motive. We can show that AI crawlers were disallowed while Googlebot was left open. We cannot show who decided it, or whether a website vendor did it on the practice's behalf without being asked.

What words do vendors use for this, and what does each one have to specify?

Most of the metric names in this field are not standard. "Share of AI voice", "AI visibility score", "citation count", "benchmark" — none of these has an agreed definition, so the name on a dashboard tells you almost nothing on its own. Whatever it is called, four things have to sit next to it before it means anything: the denominator, the engine, the date, and how many runs it came from. A number missing any of the four is a shape, not a measurement.

This is not a list of correct metrics, because we are not in a position to hand one down. It is a list of questions to put to any number, including ours. The same name can measure two different things at two vendors, and a buyer comparing them is then comparing nothing.

The right-hand column is what has to be published for the term to be checkable, not a claim about what any particular vendor means by it.
Term you will seeWhat it has to specify before it means anything
Share of AI voiceShare of what, exactly: of questions asked, of answers returned, of named brands within those answers? Plus the question set, the engine, the date and the run count. Without the denominator this is a percentage of an unstated whole
AI visibility scoreThe inputs, and whether the weighting between them is published. A score built from undisclosed weights cannot be compared to another score built from different undisclosed weights
Citation countWhether it counts appearing as a cited link, being named in the answer text, or both — and out of how many questions. These are different events and only one of them puts your name in front of a patient
BenchmarkWho is in the comparison group, how they were selected, and when they were measured. A benchmark against an unnamed peer set is a number without a population
VolatilityHow many runs, how far apart, on which engine. Two runs can show that an answer is unstable; they cannot produce a rate of change, and a rate presented from two runs is invented

This page already defines the three states we sort a result into, further up. Those definitions are the reason we can be specific here: a term like "citation count" is ambiguous precisely because it collapses two of those states into one number.

The gap is not hypothetical. Cube publishes metrics called "Share of AI Voice" and "AI Visibility Score" by name in a dental case study, and that page states no sample size, no measurement window, and no account of how the score is calculated; the practice it describes is not named either. The figures themselves are not reproduced here, because a number nobody can check does not become checkable by appearing on a second website. What each firm in this market publishes about itself, column by column, is in dental AI visibility agencies compared.

Source: cubehq.ai dental case study, read 17 August 2026.

Two of our own numbers show why the denominator rule is not pedantry. In the June 2026 study, ChatGPT returned a usable reading for 415 practices and Perplexity for 297 out of the same 527 — so "the engine named 3" and "the engine named 18" sit on different bases, cannot be averaged, and produce no single engine rate. Why answers move between runs at all, and what that does to any metric built on a single reading, is on how often do AI answers change. The underlying counts are downloadable from the Citevio data library, which is the only real test of a metric: whether a stranger can recount it.

For a location-by-location application of these denominator rules, see AI visibility for multi-location cosmetic dental groups. It keeps group-brand mentions, office mentions and cited URLs as separate fields.

See where your practice stands today

The fastest way to understand this method is to run it on your own clinic. The free checker below reads your website and scores its AI-readiness in seconds, using the same categories described above. No email and no call.

Want the manual version first? See how to check what ChatGPT says about your practice, or the deeper AI visibility checker for dentists.

Frequently asked questions

What is an AI visibility score and how is it calculated?

An AI visibility score is a 0 to 100 read of how ready a practice is to be recommended by AI assistants. Our score checks five categories: whether AI crawlers can reach the site, whether its structured data lets engines read the practice correctly, whether the site is in Bing's index, the strength of reputation signals like reviews and consistent details, and what ChatGPT, Perplexity, Gemini and Google AI Overviews actually say when asked a patient-style question. The framework is public. The weighting behind the number is ours, so the score stays hard to game.

Which AI engines do you track?

ChatGPT, Perplexity, Gemini and Google AI Overviews, because those are the assistants patients use to find a dentist. Each builds its answers differently: ChatGPT reads Bing's web index, Perplexity reads the live web with visible citations, and Gemini leans on Google's index and Google Business Profile. Where a scan checks Gemini by hand rather than through the automated pipeline, the report says so.

How often do you re-scan?

Every report shows the date its scan was run, so you always know how fresh the data is. Tracked cities are re-scanned periodically rather than on a fixed public calendar, and because a single scan is a point-in-time picture, we run each query several times within a scan. When a city is re-scanned, the report is updated and dated again. Clients see this on a running basis rather than waiting for the next public report — see how we track AI visibility for a client practice.

Why do AI answers change between runs?

AI answers are rebuilt each time from a shifting set of sources, so the same question can name one practice today and a different one next week. According to a 2025 Profound analysis of about 80,000 prompts per engine, roughly 40 to 60% of the domains an engine cites for a question change within a month. That is why we judge visibility by how often a practice appears across several runs, not by a single result. You can see this reading applied in our Charlotte dentists report.

What counts as a citation in your measurement — does a search for my practice by name count?

No. A mention only counts if it comes back for a commercial-intent question — one where a patient is choosing a treatment, a place, a price or a timeline. If an engine names your practice because someone typed your name, or quotes you inside a general-knowledge answer, we record that but don't count it: it shows the engine knows you exist, not that it puts you forward when a patient is deciding.

What five parts make up the question set you track for a practice?

A fixed core carried over from our published scans, questions built around the treatments the practice actually offers, decision-stage questions, five questions that name the practice directly to check what engines get factually wrong about it, and, on the top plan, competitor comparisons. The five name-check questions are tracked, but scored separately from the rest of the set.

If an engine names my practice only because someone searched for it by name, does that count toward the citation guarantee?

No. The named-practice questions in your set are tracked separately from the ones the citation guarantee is judged on. Being named in an answer to your own practice's name shows the engine knows you exist — it doesn't show the engine puts you forward to a patient who hasn't typed your name, which is what the guarantee actually measures.

Why does the crawler-access category in your score draw on two separate studies instead of one?

Crawler access is the one score category we've measured at national scale, and one dataset wasn't enough to back it. One study reads robots.txt across 6,497 US dental practice websites; the other requests the homepage under seven different crawler identities to check whether the server actually honors what that file promises. Both feed the same category because a robots.txt file that allows crawlers and a server that actually lets them through are two different things to verify.

Does a higher AI visibility score guarantee my practice will be recommended by an engine?

No. We measure, we don't promise: a higher score lines up with getting recommended more often, but no honest method can guarantee a spot in an answer that no agency controls. The score is a read of how ready a practice is to be recommended, not a forecast of what a specific engine will say on a specific day.

That distinction also defines the only financial commitment attached to our published plans. Citevio publishes its pricing: Visibility is $1,400/mo plus $1,900 setup, Authority is $2,900/mo plus $3,500 setup, and Dominance is $5,500/mo plus $7,500 setup. Citevio does not guarantee placement, but it does guarantee the fee: if at least two of the four engines have not named the practice in one of the commercial-intent questions agreed at onboarding within 45 days of confirmed setup, the setup fee is refunded in full.

What does a term like ‘share of AI voice’ or ‘benchmark’ need to specify before it actually means anything?

Four things: the denominator, the engine, the date and how many runs it came from. A number missing any of the four is a shape, not a measurement — the same term name can describe two different things at two vendors, and a score built from undisclosed weights can't be compared to another built on different undisclosed weights. We keep a short table on this page of what terms like ‘share of AI voice,’ ‘citation count’ and ‘benchmark’ each need to specify before they're checkable.

Why can't ChatGPT's and Perplexity's ‘named’ counts from the same scan just be averaged into one visibility rate?

Because they don't share a denominator. In the June 2026 study, ChatGPT returned a usable reading for 415 of the 527 practices scanned, and Perplexity for 297 of the same 527 — different bases from the start. A raw count of how many each engine named would be riding on those different denominators, so averaging them into one rate would hide exactly the thing a reader needs to check before trusting it.