Google reviews · June 2026 scan

Do Google reviews affect ChatGPT recommendations for dentists?

We compared Google review counts and star ratings against the practices ChatGPT and Perplexity actually named in our June 2026 scan. Here is what the review data separated, what it did not, and why neither result proves causation.

Not in one uniform way. In Citevio's June 2026 scan, Google star ratings did not distinguish practices named from those not named by ChatGPT or Perplexity. Review count showed no measurable relationship on ChatGPT, where only three practices were named, but was higher among practices Perplexity named. That is an observed association in a small, uneven sample—not proof that reviews cause AI recommendations.

Review advice for dental practices is usually written as if the relationship were settled: collect more reviews, hold a high rating, and the assistants will start naming you. We tested that on our own scan records instead of repeating it. This page reports what the review data did and did not separate, on each engine, with the sample sizes in view.

What did the scan find about reviews and AI naming?

Star rating separated nothing on either engine. Review count separated nothing on ChatGPT, where only 3 practices were named, but named practices had a higher review-count median on Perplexity.

“Named” means a practice had a VISIBLE result in the engine field of the scan record. Each row is one deduplicated dental-practice website, not necessarily one independent clinic location. We compared the Google rating and review count in that row with named versus not-named status.

Citevio E.1 analysis, 14 August 2026, over the June 2026 raw scan records. AUC is the probability that a named row has a higher value than a not-named row, ties shared. p-values are two-sided Mann–Whitney tests, exploratory, not corrected for multiple testing.
Engine and measureValid review + engine rowsNamed practicesNamed medianNot-named medianAssociation result
ChatGPT review count4113296263.5AUC 0.594; two-sided p=0.576; no measurable separation
ChatGPT star rating41134.94.9AUC 0.442; two-sided p=0.723; no measurable separation
Perplexity review count29718578239AUC 0.704; two-sided p=0.0037; positive exploratory association
Perplexity star rating297184.94.9AUC 0.586; two-sided p=0.199; no measurable separation

The star-rating result is the consistent part: the two groups had the same 4.9 median on both engines. The review-count result is not uniform. With only three ChatGPT names, the scan cannot establish a useful ChatGPT relationship. In Perplexity, named practices had a higher review-count median.

Does the Perplexity result mean that collecting reviews earns a recommendation?

No. The Perplexity finding is an association in a cross-sectional scan, not a before-and-after test of collecting reviews.

The result remained in the subset with stored, query-level evidence that agreed with the overall status: 15 named practices among 265 valid rows, with review-count medians of 573 versus 239, AUC 0.730 and two-sided p=0.0027. That sensitivity check makes the signal worth reporting, but it is not an independent replication and does not show why the practices were named.

City, specialty, brand recognition, website information, query intent and other unmeasured factors may relate both to review count and to an engine's answer. The scan did not assign practices to collect reviews or hold those factors constant. A practice should not treat this result as a forecast that a review campaign will change Perplexity or ChatGPT.

Did star ratings separate practices named by either engine?

No. In this scan, the named and not-named groups both had a 4.9 median Google star rating for ChatGPT and for Perplexity.

This says only that star rating did not separate the groups here. It does not mean that service quality, accurate profiles or genuine patient feedback do not matter. Google says review count and positive ratings can help local ranking, but Google local ranking and an AI answer are different outcomes. This analysis does not establish that improving a rating changes an AI recommendation.

Do the common 50-review and 4.7-star claims hold up in this scan?

No. Tests of the two published claims did not support either as a verified rule for ChatGPT or Perplexity in this sample.

We tested “50 or more reviews,” “4.7 stars or more,” and the two together as binary group comparisons. Fisher's exact tests did not support any of the three claims: ChatGPT p=1.000 for each comparison; Perplexity p=0.142, p=0.490 and p=0.085 respectively. Those figures are tests of claims that already circulate online, not targets we recommend.

Same E.1 analysis, 14 August 2026. Fisher exact tests, two-sided. “Not supported” means the test did not separate the groups; it is not proof that the values have no effect.
EngineClaim testedNamed among practices meeting the claimNamed among other practicesFisher two-sided pReading
ChatGPT50 or more reviews3/3540/571.000Not supported in this sample
ChatGPT4.7 stars or more3/3500/611.000Not supported in this sample
Perplexity50 or more reviews18/2600/370.142Not supported in this sample
Perplexity4.7 stars or more17/2531/440.490Not supported in this sample

Most practices meeting either claim were still not named. The tests also cannot prove that the values have no effect; they show that these two sharp claims were not supported as engine rules in this dataset.

Where can you compare your practice's review count?

Use the separate dental practice Google review benchmark for the June 2026 distribution, percentiles and measurement method. This page answers a different question: whether review data separated named from not-named practices.

Keeping the questions separate matters. A benchmark describes what was observed in a sample. It does not tell a practice what number will produce an AI recommendation.

Why are the 523 review profiles and 64 answers not one dataset?

They measure different things and must not be combined. The 523 profiles are June 2026 Google review readings for the market scan; the 64 answers are Citevio's August 2026 agency-selection measurement.

The market scan produced 523 valid Google rating and review-count profiles after production de-duplication. The 64-answer set is 16 agency-selection questions across two engines and two runs, used to measure Citevio's own visibility. It is not a clinic-level answer set, so it is not in the numerator, denominator or correlation on this page.

What are the limits of this analysis?

The positive groups are small, engine coverage is uneven, and the scan is observational. The result is useful for rejecting overconfident claims, not for proving a review tactic.

  • Small positive groups. ChatGPT had only 3 named practices in 411 valid review-and-engine rows.
  • Uneven engine coverage. Perplexity had 18 named practices in 297 rows; only 15 named cases were in the query-level evidence subset.
  • Systematic API errors in the raw scan. 112 ChatGPT API_ERROR_429 records and 230 Perplexity API_ERROR_401 records among the 527 included sites.
  • Partial evidence retention. Historical Perplexity records do not retain answer bodies, and only part of the historical scan has query-level evidence.
  • No repeated runs. The records have no repeated-run field, so model variation was not measured.

For the wider diagnostic process, see why a dental practice may not appear in ChatGPT, use the manual ChatGPT visibility check, and read how Citevio measures AI visibility. You can also inspect the open dental market research data.

Should a practice buy Google reviews to improve AI visibility?

No. This analysis does not provide a reason to buy or manipulate reviews, and Google and FTC rules address fake reviews and incentives conditioned on a particular sentiment.

Google's user-generated content policy and the FTC's Consumer Reviews and Testimonials Rule Q&A set the relevant rules. Ask for genuine feedback through a neutral process. This page is not legal advice.

What the review policy actually forbids — including the part most practices get wrong

Almost every practice knows buying reviews is against the rules. The part that surprises people is that asking only your happy patients is against them too. Google's policy names it directly — "selectively solicit positive reviews from customers" — and puts it in the same list as paying for reviews, not in a softer one. A filtering step that sends satisfied patients to Google and unsatisfied ones to a private form is that practice, whatever the software calls it.

The policy is worth reading in its own words rather than through a summary, because the wording is more specific than the paraphrases in circulation. The standard it sets first:

Google's Maps user-generated content policy

"Contributions to Google Maps should reflect a genuine experience at a place or business."

And three things a business is told not to do, quoted exactly:

  • "Offer incentives – such as payment, discounts, free goods and/or services - in exchange for posting any review or revision or removal of a negative review."
  • "Discourage or prohibit negative reviews, or selectively solicit positive reviews from customers"
  • "merchants should not require or pressure users to leave ratings or write reviews while on the premises"

Read the second one against how review software is usually sold. "Send the invite only to patients who rated us well in the internal survey" is selective solicitation of positive reviews. So is a two-step flow whose first step is a satisfaction question and whose branch decides who reaches the public form. The policy does not carve out an exception for doing it politely or automatically.

The third one catches something practices do without any software at all — the tablet at the front desk while the patient is still in the building. Google addresses that case specifically.

On what happens if a profile falls foul of this, Google states the range rather than a schedule: "we may take actions that will range from suspending the account privileges to account termination." Whose profile is at risk is the part worth being clear about — it is the practice's, not the vendor's. When any of this would actually be enforced is not something we can tell you, and nothing here is a legal assessment; it is a reading of a published platform policy.

Source: Google Maps user-generated content policy, read 18 August 2026.

What a compliant process looks like instead — everyone asked, nobody screened, no incentive, no pressure — is set out with the review distribution data in how many Google reviews a dental practice has. The underlying counts are in the Citevio data library.

We measured Google review profiles. What about Yelp and the rest?

We did not measure them, and that is the answer rather than a preamble to one. This study read Google Business Profile review counts and ratings for 523 practices. How much Yelp, Healthgrades or Zocdoc are drawn on by AI engines when they answer a patient's question was not measured here, so we cannot tell you whether investing in those surfaces moves anything.

The scope limit runs through the directory side too. The only two directories checked in this scan were Bing Places, where a usable reading came back for 291 practices out of 527, and Foursquare, where it came back for 146 — the remainder returned measurement errors on those fields and were excluded rather than counted as absent. Yelp was not in the scan at all.

So there are two claims this page will not make in either direction: that AI engines read Yelp, and that they do not. There is no comparison figure between review surfaces here, and any ratio of one against another would be manufactured. Not having measured something is also not evidence that it does not matter — Yelp may well carry weight we have not looked for.

Source: Citevio dental practice web anatomy study, 6–26 June 2026. Review profiles n = 523 of 527 (four practices had no readable rating data); Bing Places n = 291; Foursquare n = 146.

What is and is not inside the measurement is listed on how we measure AI visibility, and the raw fields are downloadable from the Citevio data library for anyone who wants to check the scope for themselves.

Reviews are one of several signals an engine can read about a practice. The checker below reads the ones on your own site in under two minutes.

Prefer to question the engines by hand? See how to check what ChatGPT says about your practice.

How we measured this

We ran the production de-duplication and dental-practice inclusion rules over antigravity/teshis_*.json, then used the Google Business Profile review fields and each engine's recorded VISIBLE or NOT VISIBLE status. The unit is a deduplicated website. The E.1 analysis was rerun on 14 August 2026 with the read-only script antigravity/e1_review_visibility_analysis_sol_20260814.py.

The scan included 527 dental-practice websites after filtering; 523 had valid rating and review-count readings. AUC here is the probability that a named row has a higher value than a not-named row, with ties shared. The p-values are two-sided Mann–Whitney tests for the continuous comparisons and Fisher tests for the published binary claims. They are exploratory and are not corrections for multiple testing.

Sources: Citevio's 14 August 2026 E.1 read-only analysis of June 2026 raw scan records; Citevio open dental market research; Google Business Profile Help; Google Maps user-generated content policy; FTC Consumer Reviews and Testimonials Rule Q&A.

Citevio's dental market research is published openly under CC-BY-4.0 with a DOI, so the numbers can be checked and reused; client data is never published.

The research package used for this review analysis is permanently referenced by DOI 10.5281/zenodo.23143214.

Common questions

Do Google reviews affect ChatGPT recommendations for dentists?

The ChatGPT part of this scan found no measurable review-count or star-rating separation, but only 3 of 411 valid rows were named. That small positive group cannot support a causal claim or a strong negative conclusion.

Do Google reviews affect Perplexity recommendations for dentists?

Named practices had higher review counts in the Perplexity rows: a 578 median versus 239, AUC 0.704 and two-sided p=0.0037. That is exploratory association, not evidence that gaining reviews causes Perplexity to name a practice.

What Google star rating does a practice need for ChatGPT?

This analysis found no verified rating requirement. The named and not-named groups both had a 4.9 median rating in the ChatGPT rows, and the scan did not test a rating intervention.

How many Google reviews does a dental practice need?

This page does not set a target. See the Google review benchmark for dental practices for the sample distribution; a distribution is not a promise about an AI answer.

Does buying Google reviews help AI visibility?

Do not buy or manipulate reviews. The study does not show that it changes an AI answer, while Google policy and FTC rules create a separate compliance risk.

Does having at least 50 Google reviews get a practice named by ChatGPT or Perplexity?

Not on its own. In our June 2026 scan, ChatGPT named 3 of 354 practices with 50 or more reviews and 0 of 57 with fewer (Fisher's exact p=1.000) — not supported. Perplexity named 18 of 260 practices with 50 or more reviews and 0 of 37 with fewer (p=0.142) — also not supported in this sample.

Does a 4.7-star rating or higher get a practice named by ChatGPT or Perplexity?

Not measurably. ChatGPT named 3 of 350 practices at 4.7 stars or higher and 0 of 61 below that mark (p=1.000). Perplexity named 17 of 253 practices at 4.7 stars or higher and 1 of 44 below it (p=0.490). Neither result supports a 4.7-star cutoff as a rule either engine follows.

Does the Perplexity review-count association hold up in a stricter subset of the data?

It stays similar, but it isn't an independent replication. In the smaller subset with stored, query-level evidence — 15 named practices among 265 valid rows — review-count medians were 573 versus 239, with AUC 0.730 and two-sided p=0.0027, close to the full-sample result of 578 versus 239 and AUC 0.704.

Could something else, like city or brand recognition, explain the Perplexity result instead of reviews?

Possibly, and we haven't ruled it out. City, specialty, brand recognition, website information, query intent and other unmeasured factors may relate to both review count and to whether Perplexity names a practice. The scan didn't hold those factors constant, so this result shouldn't be read as a forecast that a review campaign will change what Perplexity says.

How much of this scan was lost to API errors instead of a real reading?

A measurable share. Among the 527 dental-practice sites in the scan, the raw records show 112 ChatGPT API_ERROR_429 responses and 230 Perplexity API_ERROR_401 responses — sites that returned an error rather than a usable reading.

Was each practice's ChatGPT or Perplexity answer checked more than once in this analysis?

No. The records behind this specific review analysis have no repeated-run field, so run-to-run variation wasn't measured here. That's narrower than the way Citevio tracks questions in a client's set, where every question goes to all four engines across at least three separate runs because the same query can return a different answer within the same hour.

Were directory listings like Bing Places or Foursquare included in this review analysis?

Only as a separate, smaller check, not part of the review-count comparison. A usable Bing Places reading came back for 291 of the 527 practices, and Foursquare for 146; the rest returned measurement errors and were excluded rather than counted as missing. Yelp wasn't in the scan at all.