There is no single answer, and the size of the disagreement is the finding. Across the 64 answers we tracked 31 agency names. Only 20 of them survived a second identical run on the same engine, and just one name repeated in both runs of both engines. On the 8 agency-hiring questions, ChatGPT read 42 distinct domains across its two runs and only 12 of them in both, an overlap of 29%.
How was this measured?
16 buying questions, 2 engines, 2 runs each, 64 answers in total, collected 7 and 8 August 2026. All 64 calls returned an answer. The questions, engines, runs, dates, search-call counts and source domains are published as a CSV so every number below can be recounted.
| Design | Value |
|---|---|
| Questions | 16, all written as a practice owner would type them |
| Engines | 2: ChatGPT and Perplexity |
| Runs per engine | 2, identical wording |
| Answers analysed | 64, all returned successfully |
| Dates | 7 and 8 August 2026 |
| ChatGPT answers where a web search ran | 30 of 32, across 66 search calls |
| Agency names tracked | 31, from a list we compiled by hand out of these answers. It is a tracked set, not a census |
| Not measured | Gemini and Google AI Overviews |
Two answers never came from a search on ChatGPT: both runs of the veneers and smile-makeover question. The engine answered from what it already held rather than from the live web. We report those two answers in the totals and flag them here, because an answer written without a search is a different kind of evidence than one written after reading eight pages.
How an agency name was counted
By hand, and that is the step in this study that carries the most judgment. We read the 64 answers, wrote down the provider names we saw, normalised the spelling variants of each into one entry, and then searched all 64 answers for those 31 entries. Source URLs were stripped from the text first, so a domain appearing in a citation does not count as a mention. Monitoring software sold as a subscription rather than an agency service was tracked separately and is not in the 31.
That means every count on this page is a count over a list we chose, not over everything the engines said. Other provider names do appear in these answers that the list does not include, so 31, 20 and 11 are floors rather than totals. The counts within the tracked list are exact and can be recounted from the CSV; the list itself is the part you should not treat as complete.
The method we use for client measurement is the same one used here, and it is written out in how we measure AI visibility.
Why were the first ChatGPT results excluded?
Quality control on the instrument, before a single figure was published. A pilot run used gpt-4o, whose knowledge cutoff is 2023. It triggered a web search on only 1 of 8 questions, and one answer said so in plain text. That run measured the settings rather than the engine, so it was discarded and is reported here rather than left unexplained.
The answer began: "As of my last update in 2023, I can't provide specific future predictions or the best agencies in 2026, but I can guide you on what to look for when choosing a GEO (Geographic Optimization) agency for dental practices." It even expanded GEO as "Geographic Optimization", which is not what the term means in this field. Both are the signature of a model answering from memory instead of reading the live web — which is precisely what the check exists to catch. After the model was corrected, searches ran on 30 of 32 answers instead of 1 of 8, and only those runs are in the 64.
An engine that answers from memory and an engine that answers after reading eight pages produce different data, and only the second is worth publishing. Every ChatGPT figure on this page comes from the corrected configuration. Readings taken under the old setting are not comparable to these and are not mixed into them. A study that never mentions a discarded run usually means nobody checked, not that nothing broke.
Which agencies came back in both runs?
20 of the 31 tracked names. These are the agencies each engine named in the answer text in both of its runs, so they survived being asked twice. 12 repeated on ChatGPT, 9 on Perplexity, and one name, ours, repeated on both engines. Repeating is not a ranking and not an endorsement.
| Engine | Named in both of its runs |
|---|---|
| ChatGPT (12) | Atelo Group, Cause Engine Marketing, Citevio, DentalScapes, Firegang Dental Marketing, Great Dental Websites, Kruxel, Ortho Marketing, Pleiades Consultancy, Rank & Rejuvenate, Vigorant, Vydhai Dental |
| Perplexity (9) | Avante Visibility, Citevio, Dental GEO, Halcy AI, Harris & Ward, KAL Dental, Rambunctious AI, Specialty Dental SEO, UltraScout |
| Both engines, both runs | Citevio only |
The two engines barely agree with each other. Of the 20 names above, exactly one appears on both lists. That is worth holding onto before anyone quotes "AI recommends X" as though the machines had reached a consensus: on this question set they produced two almost entirely separate shortlists.
Which agencies appeared in only one run?
11 of the 31 tracked names. None of them was named in both runs of the same engine: Brux Dental Marketing, Cardinal Digital Marketing, Clarity Digital, Delmain, Dentainment, First Page Sage, Growth Honcho, My Social Practice, Practice Cafe, Studio 8E8 and Titan Web Agency. Eight appeared in a single answer. Brux Dental Marketing, Delmain and Studio 8E8 each appeared in two answers, but never in both runs of one engine.
This list is kept separate on purpose. A single-run appearance is a real observation, and it is also the exact kind of observation that turns into a false claim when it gets rounded up to "the engines recommend these agencies". They were named in one run. On the second identical run of that engine, they were not.
None of this says anything about the quality of the 11. Some of them may well be stronger operators than names on the repeated list. What the data shows is that on this question set, at this moment, their presence in an AI answer was not reproducible.
How much does an engine's answer move between two identical runs?
Far more than most people assume, and the two engines behave very differently. On the 8 agency-hiring questions, ChatGPT read 26 domains in the first run and 28 in the second, with only 12 appearing in both. That is an overlap of 29% of everything it read. Perplexity held 68% on the same questions. Across all 16 questions the figures were 40% and 77%.
| Engine and question set | Run 1 | Run 2 | In both | Overlap |
|---|---|---|---|---|
| ChatGPT, 8 agency questions | 26 domains | 28 domains | 12 | 29% |
| ChatGPT, all 16 questions | 33 domains | 34 domains | 19 | 40% |
| Perplexity, 8 agency questions | 99 domains | 104 domains | 82 | 68% |
| Perplexity, all 16 questions | 174 domains | 175 domains | 152 | 77% |
This matches what others have measured over longer windows. Profound found that roughly 40 to 60% of the domains cited in AI answers changed within a month for identical prompts. Our numbers say the churn starts much earlier than a month: on ChatGPT, 14 of the 26 domains read on the first day were not read again the next day.
Source: Profound, July 2025; Citevio study, August 2026.
The practical consequence for a practice owner is unglamorous. A screenshot of ChatGPT naming an agency, or naming your practice, is a photograph of one moment. It is evidence that something is possible, not evidence of a position you hold. This is why our own reporting counts a result only when it repeats, and why a single good screenshot should not move a hiring decision.
Where did Citevio land in its own study?
One name of the 31 tracked repeated in both runs of both engines, and it was ours. Citevio was cited as a source in 12 of 64 answers and named in the answer text in 8 of 64. Those appearances cluster by question type rather than spreading evenly: the niche, tool and pricing questions returned Citevio, the general hiring questions returned broader names.
| Measure | ChatGPT | Perplexity | Total |
|---|---|---|---|
| Cited as a source | 2 of 32 | 10 of 32 | 12 of 64 |
| Named in the answer text | 2 of 32 | 6 of 32 | 8 of 64 |
| Questions where Citevio appeared at all | 5 of 16, all of them niche, tool or pricing questions | ||
| Named in both runs of the engine | Yes | Yes | Only name of the 31 on both engines |
The split is clean. Every question Citevio appeared on contained either a narrow niche term, cosmetic dentistry, Invisalign or veneers, or referred to a free checking tool, or asked about published pricing. Every question it missed was a general one: best agency, how do I choose an agency, what are the red flags, which agency tracks visibility in a dashboard, how much should a practice spend, do reviews affect ChatGPT recommendations, how do I appear in AI Overviews.
Read plainly, that is category matching rather than reputation: an engine reaches for a narrow name when the question carries the niche, and for broad names when the question is general. Reproducibility is the other half of the reading, and it moves the other way. Of the 31 tracked names, 11 could not survive a second identical run of the same engine, and one appeared in both runs of both engines. A narrow footprint that repeats is a different thing from a wide one that does not, and this study measures both. Both halves are published, misses included, because a study that flattered its author would not be worth reading.
Two related things we can show, rather than assert: our published pricing for GEO work, and the open dataset behind our other research, 527 dental practices across 7 US metros, released under CC-BY-4.0 with a DOI. What each agency in this study publishes about its own service is set out side by side in dental AI visibility agencies compared.
What can this study not tell you?
Whether any of these agencies is good. It measures what two engines said on two days about 16 questions we wrote ourselves. It contains no client interviews, no verified results, no audit of anyone's work, and no measurement of Gemini or Google AI Overviews.
- It is not a ranking. Appearing more often means an engine found more readable material about a firm, which is a fact about web presence rather than about the work.
- The question set is ours. 16 questions written to resemble how a practice owner searches. A different 16 would produce a different list.
- The name list is ours too, and it is not complete. Agency mentions were counted against 31 names we compiled by hand from the answers. Other provider names appear in the same answers that this list does not track, so 31, 20 and 11 are floors, not totals. A name missing from the list is missing from every count on this page.
- Two engines, not four. Gemini and Google AI Overviews were not measured here, so nothing on this page describes them.
- Two runs is a floor, not a ceiling. It is enough to separate repeatable from one-off, and not enough to call anything a trend.
- We are in it. Citevio designed the questions and ran the calls, and Citevio appears in the results. The derived data is published so anyone can recount every figure here.
- We publish counts, not the models' opinions. The engines also wrote judgements about several named companies. Those are unverified statements generated by a model, not findings of ours, so they are not reproduced.
On that last point: the raw answer texts are deliberately not published. They contain characterisations of third-party firms that we did not investigate and cannot stand behind. Republishing them would spread claims we have not checked, with our name attached. The derived data carries every measurement without carrying the models' commentary.
One piece of context on why any of these firms appear at all. Ahrefs, looking at brand mentions across roughly 75,000 brands, found mentions of a brand on other sites tracking Google AI Overview appearances at r=0.664, against 0.218 for links pointing at the site. Being written about elsewhere is a large part of what makes a company visible to an engine, which is also why a study like this measures the shape of the web more than the quality of the work.
Source: Ahrefs, May 2025. That analysis covers Google AI Overviews specifically, not ChatGPT or Perplexity.
For scale, SOCi's 2026 local visibility index reported that ChatGPT recommended 1.2% of the locations it studied, Perplexity 7.4% and Gemini 11%. Being named at all is rare on these surfaces, for practices and for agencies alike.
Source: SOCi 2026 Local Visibility Index, reported by Search Engine Land, January 2026.
You can run a version of this on your own practice name in under two minutes.
Prefer a deeper look? The full AI visibility checker for dentists tests what ChatGPT, Perplexity, Gemini and Google AI Overviews actually say about you.
Common questions
Does being named by ChatGPT mean an agency is good?
No. This study measured what two engines said on two days, nothing else. We did not review anyone's work, speak to their clients or verify a single performance claim. An engine names an agency because it found and read pages about it, which is a fact about web presence, not about results.
Why were the first ChatGPT results excluded?
Quality control on the instrument, before any figure was published. A pilot run used gpt-4o, a model with a 2023 knowledge cutoff, and it triggered a web search on only 1 of 8 questions. One answer said in plain text that as of its last update in 2023 it could not name the best agencies in 2026, which is the signature of a model answering from memory rather than from the live web. That run measured the settings rather than the engine, so it was discarded and is reported here rather than left unexplained. After the model was corrected, searches ran on 30 of 32 answers, and those are the runs this study reports.
How much does an AI engine's agency shortlist change between two identical runs?
A lot on ChatGPT. Across the 8 agency-hiring questions, ChatGPT read 26 non-platform domains in run one and 28 in run two, and only 12 appeared in both, an overlap of 29%. Perplexity was steadier at 68% for the same questions. Any list built from a single run should be read as one sample, not a ranking.
Citevio ran this study. How did Citevio do in it?
Citevio was the only one of the 31 tracked agency names that repeated in both runs of both engines. It was cited as a source in 12 of 64 answers and named in the answer text in 8 of 64. Those appearances cluster by question type rather than spreading evenly: the niche, tool and pricing questions returned Citevio, and the general hiring questions returned broader names.
Can I check these numbers myself?
Yes. Two CSV files are published under CC-BY-4.0: one row per answer with the question, engine, run, date, search calls and source domains, and one row per agency with its counts. Every number on this page can be recounted from them. One caveat you should carry into the recount: agency mentions were matched against a list of 31 names we compiled by hand, so the counts are complete for those 31 and not for every provider the engines named. The full answer texts are not published, because they contain the models' own unverified statements about named companies.
Download: answer-level data, 64 rows (CSV) · agency-level counts, 31 tracked names (CSV). Both CC-BY-4.0.