All posts
Research56 min read

Asking Every Patient Pays. The Score Is Just the Trigger.

Twenty-seven sources on what patient feedback is actually worth to a private practice: 5–9% of revenue per displayed star, a 5.8-fold malpractice-risk gradient in complaint counts, and why asking everyone is both the compliant design and the effective one.

Tom North·Founder, applaud·

Bottom line. Asking every patient how their visit went pays for itself three separate ways, and none of them depend on the score being a good score. The number is a trigger. The sentence the patient writes underneath it is the finding.

The short version

Six things this review establishes, each of them independent of whether any particular metric is valid:

  1. Displayed ratings move money, and they move it most for small practices.Two quasi-experimental studies put the causal effect of a displayed star at 5–9% of revenue, concentrated entirely in independent, uncredentialed, low-review-volume businesses (Luca, 2016; Anderson and Magruder, 2012). The practices that judge themselves too small for this to matter are the population in which the effect is measured.
  2. The count of unhappy patients predicts your risk.Unsolicited complaint volume predicted malpractice claims after adjusting for clinical volume, with a 5.8-fold gradient between physicians holding two or more lawsuits and those holding none (Hickson et al., 2002). You cannot count what you never invited.
  3. Reputation decays in weeks, and responds mainly to being asked. Recency dominates how a profile reads, and solicitation is the dominant driver of review volume (BrightLocal, 2026; Han et al., 2024). So the real choice is not between running a feedback process and running none. It is between one you run and one run for you by whichever patients were annoyed enough to post unprompted.
  4. The score is a router, not a KPI. This is the honest answer to what the eleven-point question is for. It is an adequate mechanism for triggering an action the same day, and a poor mechanism for measuring a practice. The free text carries the diagnostic content: for a given service feature, the written description measures perceived quality better than the number attached to it (Xu et al., 2021), and the comments field is independently identified as the most useful component of the survey (Adams et al., 2022).
  5. Passives are where the damage hides.Scores of 7–8 are excluded from the NPS calculation entirely, yet 37.8% of them came with negative or mixed written feedback against 6.7% of promoters — a 5.6-fold difference. Route on what the patient wrote, never on which band the number fell in.
  6. Asking everyone is both the compliant design and the effective one.HIPAA expressly permits collection, follow-up and internal analysis as a health care operation. What carries federal penalty and platform termination is the shortcut — soliciting only the patients you expect to say something nice. There is a $4.2 million enforcement precedent on the record.

What to actually do

The specification below is the minimum the evidence supports. Every element is here because a finding requires it, and it is set out in full with its citations in section 10.2.

DecisionWhat the evidence supports
Who gets askedEvery patient. No sentiment filter at any point — that is what keeps you inside the FTC's generalized-solicitation exemption and inside platform policy.
WhenOnce, 24 to 72 hours after the visit.
What you sendTwo items: an eleven-point rating, then a free-text prompt. No offer or promotion in the same message — bundling one reclassifies it as telemarketing and raises the consent bar.
What triggers a callAny score of 0–6, and any negative free text at any score, to one named person the same day.
Who gets the review linkEveryone who responds, in the same message, whatever they scored.
What you measureThree numbers monthly: detractor count (not the mean), response rate, and three to five recurring themes from the text.
What never to doAsk only happy patients. Offer anything of value for a review. Put a review screen in front of someone on the premises. Compare your score to a published benchmark.

Why NPS matters

Not because the score predicts growth — it does not, and we say so below. Because its two ingredients are the ones tied to what a practice is worth, and because a number is what makes every patient sortable on the day they are still reachable.

The downside is the larger quantity, and the score is how you find it.A panel of 19,792 hospital-year observations across 3,767 US hospitals found positive patient experience associated with increased profitability and negative experience associated — more strongly — with decreased profitability (Richter and Muhlestein, 2017). The asymmetry is the finding: a programme built to catch detractors is better matched to the evidence than one built to raise an average. Two caveats belong with it. The study reports direction without coefficients, so no magnitude is claimed here; and its hospitals operate under a reimbursement mechanism cash-pay practice does not share, so what transfers is the reputational and retention channel, not the payment one.

And its four variables are a decomposed NPS.That panel did not measure a net score. It measured the percentage who would definitely recommend, the percentage who would definitely not, the percentage rating 9–10 and the percentage rating 6 or below — promoters and detractors tested separately rather than subtracted from one another. Which is the precise shape of the argument: the ingredients carry the association with profitability. The arithmetic that nets them into a single headline number is the part that adds nothing.

Intercepting a complaint is worth roughly three times replying to one.A factual negative review depresses selection intention by about 1.22 points more than a vague one, and a public reply recovers around a third of what preventing that review would have been worth (β = +0.194 against −0.371). A complaint you catch privately never enters the denominator at all. This is also the one part of the workflow the Federal Trade Commission expressly endorses in its own guidance on the Consumer Reviews Rule.

And the risk sits in the count.Unsolicited complaint volume predicted malpractice claims after adjustment for clinical volume, with a 5.8-fold gradient across the distribution (Hickson et al., 2002). That is a count of dissatisfied patients — exactly what the detractor band gives you, and exactly what you cannot see without asking.

Two further pieces point the same way and are reported here with their limits attached. A systematic review of fifty-five studies found patient experience positively associated with patient safety and clinical effectiveness, positive associations outnumbering null findings 429 to 127, though the authors are explicit that association is not causation (Doyle, Lennox and Bell, 2013). And promoters have been measured spending 13% more than passives and 46% more than detractors — a genuine per-customer differential and a plausible mechanism, but general-industry, not healthcare, and reported at one remove from the primary source (Mecredy et al., 2018, in Sauro, 2019).

The critique is narrower than it sounds.Replacing the single recommendation item with multi-item loyalty measures increased explanatory power by only 0.1 to 1.6 percentage points of R² (Keiningham et al., 2007), and NPS correlates with the American Customer Satisfaction Index at r = 0.99, 0.72 and 0.92 (East et al., 2011). NPS is not the best available instrument. It is also not materially worse than the alternatives, and it is the one question patients already recognise, which is why it gets answered. Whether a structured signal is collected at all matters considerably more than which instrument collects it.

The claim we are not making

NPS is sold to practices on the promise that it predicts revenue growth. That claim does not survive contact with the literature: two independent replications, one using the original methodology, failed to find the relationship (Keiningham et al., 2007; Baehre et al., 2022), and in direct comparison the metric performed worse than a single global rating item (Adams et al., 2022).

We say so here because the argument for collecting patient feedback does not need it, and is weakened by being attached to it. The case above rests on complaint counts, displayed-rating economics and regulatory position — three literatures, none of which depends on the metric being valid. Whether a structured signal is collected at all matters considerably more than which instrument collects it.

The evidence, in full

What follows is the complete review: twenty-seven sources, a stated evidence-classification scheme, an adversarial verification procedure, the regulatory perimeter section by section, and an appendix listing the claims we examined and decided not to assert.

A narrative evidence review of predictive validity, reputational economics, and regulatory constraint in the United States

Working paper. Compiled 3 September 2026.

Scope: privately-owned dental, medical-aesthetic and medspa practices operating in the United States.

Abstract

Background. The Net Promoter Score (NPS) is widely marketed to private healthcare practices on the claim that it predicts revenue growth. Small practices frequently decline to collect structured patient feedback on three grounds: that their respondent volume is too low for a score to be meaningful, that they already know what their patients think, and that publicly posted reviews are an adequate substitute.

Objective. To establish what the available evidence does and does not support regarding the systematic collection of patient-experience feedback in small and medium private practices, and to identify the regulatory constraints governing its collection and use in the United States.

Methods. A five-strand parallel literature and regulatory search was conducted on 29 August 2026 across peer-reviewed databases, federal regulatory sources, and industry publications. Twenty-seven sources were retrieved and read. One hundred and thirty-five candidate claims were extracted; twenty-five were subjected to a three-vote adversarial verification procedure requiring two of three independent refutations to reject a claim. Sources were assigned to one of five evidence classes, and claims failing verification are reported rather than removed.

Results. The growth-prediction claim was not supported. Two independent replications, one using the original methodology, failed to detect a relationship between NPS and revenue growth by simple correlation (Keiningham et al., 2007; Baehre et al., 2022). NPS has not been validated for healthcare use and performed worse than a single global rating item in direct comparison (Adams et al., 2022). Three independent literatures, none contingent on NPS validity, did support systematic collection. First, the count of unsolicited patient complaints predicted malpractice risk after adjustment for clinical volume, with a 5.8-fold gradient between physicians with two or more lawsuits and those with none (Hickson et al., 2002). Second, quasi-experimental estimates identify a causal effect of displayed rating on revenue of 5–9% per star, concentrated entirely in independent, uncredentialed, low-review-volume businesses (Luca, 2016; Anderson and Magruder, 2012). Third, collection is expressly permitted under HIPAA as a health care operation, whereas the practice of soliciting reviews only from patients presumed satisfied carries documented enforcement exposure under Section 5 of the FTC Act and platform-policy termination risk.

Conclusions. The case for systematic collection does not rest on the predictive validity of any score and is weakened by being argued on that basis. The defensible position is that the numeric item functions adequately as a routing trigger and poorly as a performance measure; that free-text response carries the diagnostic content; that detractor count rather than mean score is the tractable statistic at low volume; and that universal, unfiltered solicitation is simultaneously the compliant design and the effective one.

Keywords: patient experience; Net Promoter Score; service recovery; online reviews; regulatory compliance; private practice; patient-reported experience measures

1. Introduction

1.1 Background

Patient-experience measurement entered private ambulatory practice largely through the commercial diffusion of the Net Promoter Score, introduced by Reichheld (2003) as "the one number you need to grow." The metric asks respondents to rate, on an eleven-point scale, their likelihood of recommending a provider to others; responses of 9–10 are classified as promoters, 7–8 as passives, and 0–6 as detractors, and the score is calculated as the percentage of promoters less the percentage of detractors.

Its adoption in healthcare has been substantially commercial rather than scientific. Vendors of reputation-management and patient-communication software typically position the metric on the growth-prediction claim, and practice owners consequently evaluate the decision to collect feedback as a question about whether that claim is true.

This framing is unhelpful in both directions. Where the claim is accepted uncritically, practices adopt a measurement programme whose central promise the literature does not support. Where it is rejected, practices frequently discard the entire activity along with the metric, forgoing benefits that do not depend on the metric's validity at all.

1.2 Objectives

This review addresses six questions:

  1. What outcomes, if any, does patient-experience measurement predict in healthcare settings, and at what effect size?
  2. What is the strongest available case against the use of NPS, and what survives it?
  3. At what respondent volume does a score become statistically interpretable, and what should practices below that threshold measure instead?
  4. What is the measured relationship between displayed ratings, review volume and patient acquisition?
  5. What is the evidentiary basis for service recovery as an intervention?
  6. What constraints do United States federal regulation and platform policy impose on the collection and use of this feedback?

A seventh question — the identification of operational parameters for instrument timing, mode and length — is addressed where evidence permits.

2. Methods

2.1 Search strategy

A five-strand parallel search was executed on 29 August 2026, with strands directed respectively at: (i) predictive-validity literature for recommendation-intention metrics; (ii) critical and methodological literature on NPS; (iii) econometric literature on online ratings and demand; (iv) service-recovery literature; and (v) United States federal regulation, agency guidance and platform policy.

Twenty-seven sources were retrieved and read in full where access permitted. Where a source was available only as an abstract owing to paywall restriction, this is stated at the point of citation and no magnitude is reported from that source.

2.2 Evidence classification

Sources are not treated as interchangeable. Each is assigned to one of five classes, and the class is stated in prose wherever it bears on the weight a claim can carry:

Table 1. Evidence classification scheme applied throughout this review.

ClassDefinition
Peer-reviewedPrimary research or systematic review published following peer review
RegulatoryCodified regulation, federal rulemaking, agency enforcement record, or platform policy
Government guidanceAgency-published guidance not subject to peer review
Vendor-publishedPublished by a commercial party selling the product under discussion
Expert panelStructured elicitation of practitioner opinion; not measurement

The distinction between the first and the last two is material to several findings reported below, and is drawn explicitly rather than left to the reader.

2.3 Verification procedure

One hundred and thirty-five candidate claims were extracted from the retrieved sources. Twenty-five claims — selected on the basis of load-bearing significance to the review's conclusions, or of wide circulation in commercial literature — were subjected to independent three-vote adversarial verification, in which a claim was rejected only where two of three independent checks refuted it. Eighteen claims were confirmed at 3–0 or 2–1; seven were refuted.

2.4 Treatment of refuted claims

Refuted claims are not silently omitted. Where a claim is both widely circulated and unsupported, it is reported together with the reason it fails, on the ground that a practitioner encountering the claim elsewhere is better served by knowing why it does not hold than by its absence. Such claims are identified in the text and collected in Appendix A.

2.5 Limitations of the method

This is a narrative review, not a systematic review conducted to PRISMA protocol. Search strands were designed to answer a practitioner-facing question rather than to achieve exhaustive coverage, and no formal risk-of-bias instrument was applied to included studies. Grey and vendor literature was deliberately included, because it constitutes the evidence base practitioners are in fact exposed to; it is labelled throughout and is not treated as equivalent to peer-reviewed work.

3. Predictive validity

3.1 The growth claim and its replications

Reichheld's (2003) original proposition — that recommendation intention is the single best predictor of firm growth — has been directly tested and was not supported.

Keiningham, Cooil, Andreassen and Aksoy (2007), publishing in the Journal of Marketing, replicated the Reichheld and Satmetrix methodology on longitudinal data from twenty-one firms and more than 15,500 customer interviews drawn from the Norwegian Customer Satisfaction Barometer over the period 2000–2005. They report "no real indication that average levels of any of the satisfaction/loyalty metrics in Table 1 are significantly correlated with the relative change in revenue." Pooled Net Promoter correlations were r = 0.40 (p = 0.20) in banking and r = −0.45 (p = 0.55) in retail gasoline.

The same study reconstructed Reichheld's own scatterplot data from The Ultimate Question for the three United States industries also tracked by the American Customer Satisfaction Index. The ACSI — which Reichheld had publicly characterised as having zero correlation with growth — produced a higher R² than Net Promoter in two of three cases. Taken with Reichheld's own acknowledgement that NPS was not the best predictor in four of his twelve industries, the authors concluded that NPS could not be classified as the best predictor in half the industries examined.

Of particular relevance to any practice treating a score as a key performance indicator: a best-subsets regression considering all eleven candidate satisfaction and loyalty metrics selected none of them. The best-fitting model under the Bayesian information criterion contained only industry fixed effects.

The most methodologically careful modern re-test reaches a partially rehabilitating verdict. Baehre, O'Dwyer, O'Malley and Lee (2022), in the Journal of the Academy of Marketing Science, analysed 193,220 NPS evaluations across seven United States sportswear brands over nineteen quarters. Their findings, ordered by relevance to ambulatory practice, were:

  • Simple correlation — the method Reichheld employed, and the method most practitioners continue to employ — detected no relationship between any NPS measure and current or future sales growth. The authors explicitly report a failure to replicate Reichheld (2003).
  • Only changes in NPS, measured across all potential customers, predicted next-quarter sales growth (β = 1.458, p < 0.01, adjusted R² = 0.350). The classic customer-loyalty formulation — current customers only, which is precisely what a practice surveying its own patients obtains — was not significant in either levels or changes.
  • Incremental explanatory power was small: inclusion of the NPS variable raised adjusted R² by 0.028.
  • Predictive validity extended one quarter only; two- and three-quarter lags produced no significant findings.
  • NPS did not outperform alternatives. Changes in brand consideration predicted growth at identical model fit (adjusted R² = 0.350 for both).

The authors conclude that NPS "can explain only a fraction of future sales growth by itself," and that it "cannot diagnose the cause of a problem, so it must be paired with follow-up questioning."

The implication for practice is direct. The one specification in which a version of the growth claim survived rigorous testing required changes in a market-wide measure over a single quarter — none of which describes what a practice's own post-visit survey produces.

3.2 Validation status in healthcare

Adams, Walpola, Schembri and Harrison (2022), in Health Expectations, conducted the only systematic review of NPS in healthcare located by this search. Five databases were searched for the period January 2005 to September 2020; 468 records were screened and twelve studies met inclusion criteria. The review's central finding is stated unambiguously: "Although NPS has been used in a wide range of service industries, it has not been validated for use in the healthcare setting."

Two features of the included literature limit its applicability to United States private practice. Ten of the twelve studies were conducted in the United Kingdom and none in the United States, despite the metric's American origin. The settings are predominantly public-sector.

The review further reports a direct comparison between NPS and a single global rating item ("How would you rate the hospital or clinic?", 0–10). The global rating demonstrated stronger associations both with quality indicators and with patient experience as measured by the Consumer Quality Index. This is the healthcare-specific analogue of the Keiningham critique: NPS did not outperform a simpler alternative. The English National Health Service subsequently moved the Friends and Family Test away from the recommendation question toward an overall rating.

Finally, and of the greatest practical significance, the review found little evidence that collection alone produces improvement. In one study of general practices, only one of forty-two clinics (2.4%) could identify an instance in which NPS results had led to a quality improvement. Four of the twelve reviewed studies found that NPS generates high data volume with limited improvement value, and the reviewers concluded that more specific questions are required alongside the metric to produce actionable insight.

3.3 Measurement properties

Rocks (2016), in a preprint whose results are reproduced in the subsequent peer-reviewed literature, establishes three properties of the NPS statistic that bear on its use at low volume.

First, the variance of the score ranges from 0 to a maximum of 1, maximised where respondents divide evenly between promoters and detractors. This is four times the maximum variance of 0.25 attaching to an ordinary binomial proportion, with the consequence that an NPS carries up to twice the standard error of a simple percentage computed at the same sample size. This is a property of the metric's construction and not of any particular dataset.

Second, the score is not identifiable from the response distribution. Many distributions produce the same score, and unlike a proportion the score cannot be recovered from its variance. A sample composed entirely of passives and a sample divided evenly between promoters and detractors both return exactly zero.

Third, the conventional confidence interval is incorrect. The standard Wald interval is systematically too narrow across the entire range tested, with coverage failing to attain 95% even at n = 300 (81.97% at n = 5; 91.07% at n = 15; 93.06% at n = 30; 94.81% at n = 300). Rocks concludes that Wald and Goodman methods "should be avoided."

Fisher and Kordupleski (2019), writing in Applied Stochastic Models in Business and Industry, advance a separate objection concerning the classification thresholds: "A company using NPS is basically not distinguishing those of its customers rating them 0 from those rating them 6, and is totally ignoring 'passives', who are rating them 7 or 8." Their argument is that the 5–6 group is plausibly convertible upward and the 7–8 group plausibly promotable, such that collapsing both discards the segment with the greatest available movement. They further observe that there is "little hard evidence that NPS is a lead indicator of business outcomes."

An evidence-class caveat is warranted here, and is not stated in the source itself. The Fisher and Kordupleski article is peer-reviewed but appears in the journal's Practitioner's Corner as a critical review presenting no new empirical data, and the second author originated the Customer Value Management method that the article advocates throughout — a competing-methodology interest the article does not declare among its limitations.

3.4 Complaint volume and malpractice risk

The strongest evidence located by this review concerns neither NPS nor any score. It concerns counting.

Hickson, Federspiel, Pichert, Miller, Gauld-Jaeger and Bost (2002), in JAMA, report a retrospective longitudinal cohort of 645 general and specialist physicians in a large United States medical group, covering January 1992 to March 1998 and 2,546 physician-years.

The total number of unsolicited patient complaints significantly predicted three separate outcomes — risk-management file openings, file openings with expenditure, and lawsuits — and did so, in the authors' words, "even when data were adjusted for clinical activity," operationalised as a log transformation of relative value units expressed as a percentage of the group average. After adjustment for RVUs, specialty and sex, complaint count remained an independent predictor: Wald χ² = 27.3 (p < 0.001) for file openings, 11.3 (p < 0.001) for openings with expenditure, and 8.3 (p = 0.004) for lawsuits. Predictive concordance was 84%, 83%, 81% and 87% for file openings, openings with expenditure, lawsuits and multiple lawsuits respectively.

The gradient is steep and monotonic. Surgeons with two or more lawsuits averaged 35.1 unsolicited complaints; those with one, 16.7; those with none, 6.1 — a 5.8-fold spread, with all three pairwise comparisons significant (p < 0.001, p = 0.001, p < 0.001).

A widely repeated secondary claim about this study does not survive checking. It is frequently asserted that what distinguished high-liability physicians was communication failure — inattentive listening, unreturned calls, discourtesy — rather than clinical error. The 2002 study does not report this. The sentence generally quoted appears in the paper's Comment section and carries citations to two earlier obstetric studies from 1992 and 1994. The study's own finding is opposite in emphasis: "In the present study, the total number of patient complaints, not any particular type, predicted risk management outcomes." The single category adding incremental predictive value was care and treatment — the clinical category. The authors' subsequent synthesis with the Agency for Healthcare Research and Quality states explicitly that complaint scores based on both clinical and interpersonal failures are associated with risk-management activity.

Two further caveats attach. The 35.1 / 16.7 / 6.1 means are unadjusted, and the authors note that complaint and risk-management data were both positively correlated with clinical volume; the volume-adjusted result is the regression, not the raw means.

The distribution is extremely concentrated. A replication in urology found that 11% of urologists generated 50% of complaints across fifteen health systems (Stimson et al., 2010).

The operational implication is that the predictive signal resides in the count, not in the theme and not in the mean. A practice recording how many patients reported a problem in a given month is tracking the variable that carried predictive weight in the canonical study. A practice recording only a mean score has discarded it.

3.5 Patient experience, safety and financial performance

Doyle, Lennox and Bell (2013), in BMJ Open, systematically reviewed fifty-five studies (forty individual studies and fifteen systematic reviews) examining links between patient experience, patient safety and clinical effectiveness. They report consistent positive associations across disease areas, settings and study designs, with positive associations outnumbering null findings by 429 to 127.

Their own caveat is material and is reproduced in full: "Because associations do not entail causality, this does not necessarily prove that improvements in patient experience will cause improvements in the other two domains."

Richter and Muhlestein (2017), in Health Care Management Review, analysed a panel of 19,792 hospital-year observations drawn from 3,767 United States hospitals over 2007–2012, using CMS and HCAHPS data and generalised estimating equations against three financial outcomes. The reported finding is that "a positive patient experience is associated with increased profitability and a negative patient experience is even more strongly associated with decreased profitability."

Two features make this study more transferable to private practice than its hospital setting suggests. Its four independent variables are structurally a decomposed NPS — percentage who would definitely recommend, percentage who would definitely not recommend, percentage rating 9–10, and percentage rating 6 or below — with the promoter and detractor components tested separately rather than netted. And the asymmetry it identifies, in which the downside of poor experience exceeds the upside of good experience, is directionally informative regardless of setting.

Three limitations must be stated. The abstract reports direction only, without coefficients, effect sizes or confidence intervals, and the full text is closed-access; no magnitude can therefore be cited from this source. The authors themselves state that the experience-to-profitability link "has not been well established in the health care industry" as of publication. And the sample comprises acute-care hospitals operating under CMS value-based purchasing, a reimbursement mechanism with no analogue in cash-pay dental, medspa or aesthetic practice; what transfers is the indirect reputational and retention channel, not the payment mechanism.

3.6 Outcomes for which no evidence was located

This review sought evidence linking patient-experience measurement to dental treatment-plan acceptance, to non-attendance and cancellation rates, and to patient lifetime value in healthcare settings. A five-strand search did not surface it in the peer-reviewed literature.

The nearest available figure is general-industry and secondary. Mecredy et al. (2018), reported at one remove in a review by Sauro (2019), measured promoters spending on average 13% more than passives and 46% more than detractors, with NPS correlating to future revenue at r = 0.91. This is a genuine per-customer spend differential and a plausible mechanism, but it is not healthcare and is reported here without access to the primary source.

Practitioners presented with a treatment-acceptance uplift figure attributed to patient-experience measurement should request the underlying study. This review did not locate one.

4. Statistical constraints at low respondent volume

The most substantively correct objection raised by small practices is that their volume is insufficient for a score to carry meaning. The objection is correct about the score and incorrect about the programme.

4.1 Minimum sample requirements

Costa and Ponte (2024), in the Anais da Academia Brasileira de Ciências, derive minimum sample sizes for estimating NPS under a Bayesian approach. Table I, computed under a uniform prior α = (1,1,1) at 95% credibility, gives:

Table 2. Minimum sample sizes for NPS estimation under a Bayesian approach with uniform prior, 95% credibility (Costa and Ponte, 2024, Table I).

Target credible-interval width (NPS points)Responses required
20 (±10)164
10 (±5)655
4 (±2)4,085
2 (±1)16,034

Sample size scales with the inverse square of interval width: halving the width quadruples the responses required, and precision improves with the square root of n.

4.2 Three misreadings of the published thresholds

The figure of 655 responses circulates widely and is routinely misstated in three respects.

First, it is an average rather than a guarantee. The paper employs the Average Length Criterion, under which n = 655 bounds the mean interval width over the prior predictive distribution. Approximately half of realised intervals will be wider, and the authors' own simulation confirms that realised widths were "slightly larger" than target and coverage "slightly smaller" than 95%. A genuine worst case — an even promoter–detractor split at maximum variance — requires approximately n = 1,537.

Second, the figure is prior-dependent. At the same target width, α = (0.1, 0.1, 0.1) yields 108 and α = (5, 5, 5) yields 860, a 6.1-fold spread. There is no constant "655 responses required."

Third, the paper's frequently quoted worked example — 406 responses producing an NPS of +12.7 with a 95% interval of [+3.8, +20.6], 16.8 points wide — is described by the authors as "a hypothetical dataset on financial services in three markets" constructed "to mimic the application of the methods." It is simulated data, and it is not healthcare.

4.3 The effect of distributional shape

The last of these points has a consequence favourable to small practices that does not appear in commercial discussion of the topic. The 16.8-point interval in the worked example is a property of the split, not of the sample size; the example's distribution is near-neutral, at 33.5% detractors and 46.3% promoters. Healthcare distributions are typically nothing of the kind, being heavily weighted toward promoters.

Recomputing on the same variance structure at a distribution more representative of healthcare — 75% promoters, 20% passives, 5% detractors, giving an NPS of approximately +70 — at the same n = 406 yields an interval approximately 10.8 points wide, some 35% narrower. Because NPS variance falls as the response distribution concentrates, a practice whose patients are genuinely satisfied obtains a more precise estimate from a given number of responses than the textbook example implies.

4.4 What practices below threshold should measure

Extrapolating the square-root relationship below the table's floor gives an interval of approximately 43 points at 36 responses and approximately 36 points at 50. A practice collecting a few dozen responses monthly therefore cannot resolve score movements smaller than tens of points, and reporting month-on-month movement at that volume is not defensible.

Two remedies have support in the same literature.

The first is temporal aggregation. Costa and Ponte demonstrate sequential Bayesian updating in which carrying the first quarter's posterior into the second — 825 cumulative responses — narrowed the interval from 16.8 to 12.0 points while the point estimate moved only from +12.7 to +13.1. Trailing-twelve-month reporting is interpretable where month-on-month reporting is not.

The second is to abandon the netted score in favour of a count. Pingitore et al. (2007) observe that net scoring — promoters less detractors, discarding passives — requires substantially larger samples than mean-based scoring for comparable precision; the instability resides in the netting operation. A raw count of detractors in a given month is a statistic a small practice can own, and it is, per Hickson et al. (2002), the count rather than the average that carried the predictive signal.

Practices below the volume threshold are therefore better served by recording: the number of detractors in the period; the number of complaints resolved; recurring themes in free-text response; and, if a score is reported at all, a rolling twelve-month figure presented with its credible interval.

5. Objections from practice: knowledge and substitution

5.1 "Practice owners already know what their patients think"

Three findings bear against this proposition.

The Agency for Healthcare Research and Quality states as an operating principle that "service recovery cannot take place if the provider does not know that the member or patient is unhappy," and lists "effective systems for inviting/encouraging customers to complain" among the required components of a service-recovery programme. Complaints reaching the front desk are a self-selected subset of complaints that exist.

A caveat attaches to that source. The same AHRQ page carries the familiar assertion that "only 50 percent of unhappy customers will complain… but 96 percent will tell at least nine or ten of their friends." That sentence carries no reference number, is attributed only to unnamed "several marketing studies," and none of the page's five numbered references support it. It should be treated as legacy commercial folklore restated by a federal agency rather than as a measured effect. The same page renders a Reichheld and Sasser retention figure as a bare "377" without unit, apparently a typesetting fault; it should not be quoted.

The Hickson data provide the sharper answer. The complaint counts analysed were unsolicited — they arrived without any survey instrument. Physicians in the highest-risk group accumulated approximately 35 complaints where colleagues accumulated 6, over a period during which, presumably, none of them believed themselves to have a problem.

Finally, respondents are not the patient panel. Johnston, Hogg, Wong, Burge and Peterson (2021), in the Journal of Medical Internet Research, studied survey response across eighty-seven primary care practices and found respondents skewed older, more female, higher-income and higher-morbidity than non-respondents. A score describes respondents; clinical intuition describes patients who speak to staff; neither describes the panel.

5.2 "Publicly posted reviews are an adequate substitute"

Publicly posted reviews are an output of a feedback process rather than a substitute for one, on four grounds.

Selection. Public reviews are written by the small tail of patients sufficiently motivated to post unprompted. BrightLocal (2026), a vendor-published consumer survey, reports that 78% of consumers were asked for a review in the preceding twelve months and that 83% of those asked went on to leave one. Reviews are overwhelmingly a function of solicitation; a practice that does not solicit is observing whichever patients reached the review interface unaided.

Timing. By the point at which a complaint has become a public review, the recovery window has closed. The interventions discussed in Section 6 depend on prior notification.

Diagnosticity. Xu, Armony and Ghose (2021), in Management Science, derived the seven service-quality dimensions patients raise in physician reviews: bedside manner, diagnostic accuracy, waiting time, service time, insurance process, physician knowledge, and office environment. A star rating identifies none of them; free text does, and in their models the text-derived measure predicted physician choice better than the numeric rating of the same feature.

Regulatory. As set out in Section 7, the public review surface cannot lawfully be engineered directly. Platform policy prohibits selective solicitation of positive reviews, prohibits incentives, prohibits on-premises pressure, and prohibits requesting specific content, with enforcement escalating to account termination. The compliant route to higher review volume runs through a universal feedback process rather than around it.

6. Reputational and acquisition economics

The causal evidence in this domain is considerably stronger than that available for the predictive-validity question, because the structure of review platforms admits a natural experiment.

6.1 Quasi-experimental estimates of displayed-rating effects

Both landmark studies exploit the same discontinuity: platforms round ratings to half-stars, so that a business with a true average of 3.24 displays three stars and one at 3.26 displays three and a half. Comparison across that threshold isolates the effect of the displayed rating from that of underlying quality.

Luca (2016; originally 2011) applied regression discontinuity to Washington State Department of Revenue tax records covering all Seattle restaurants over 2003–2009. The principal findings were:

  • A one-star increase in displayed rating causes a 5–9% increase in revenue. The estimate is causal rather than correlational.
  • The effect is concentrated entirely in independent businesses. For chain-affiliated restaurants the estimate is "statistically insignificant and close to zero," the brand having already resolved quality uncertainty.
  • Review volume moderates effect size. A rating change carries approximately 50% greater revenue impact at businesses with fifty or more reviews than at those with fewer than ten.
  • Consumers respond to the rounded displayed rating rather than to the finer underlying average, so the return to feedback is threshold-driven rather than continuous.

Anderson and Magruder (2012), in the Economic Journal, applied the same design to 328 San Francisco restaurants matched to a reservation platform over July–October 2010:

  • An additional displayed half-star causes restaurants to sell out their 7 p.m. reservations 19 percentage points more often, a 49% relative increase.
  • The effect is concentrated entirely among businesses lacking an independent quality credential. For restaurants without a Michelin star or Chronicle Top 100 listing, an additional half-star reduced availability by 20–30 percentage points across all three dinner services; externally accredited restaurants showed no such gain, with the difference significant at the 1% level.
  • The effect is likewise concentrated among lower-volume businesses. Below 500 reviews, the 20–30 percentage-point effect obtains; above 500 reviews, no discontinuity is detectable at any threshold.
  • The authors' own calibration, which they flag as approximate, gives the median restaurant a 6–9% gain in customer flow per additional half-star; at $20,000 weekly revenue and a 68% margin, a 6% lift represents $816 weekly pre-tax profit against roughly $2,000 weekly median profit — an approximate 40% profit increase.

Both studies point in the same direction, and it is the opposite of the inference small practices typically draw. The rating premium is largest for the unbranded, uncredentialed, low-volume business. A single-operator practice has more at stake in its displayed rating than a multi-site group with an established name, not less.

The transferability caveat must be stated plainly. Both settings are restaurants. Applying a 5–9%-per-star coefficient to a dental or aesthetic practice is analogy rather than measurement. What transfers with greater confidence is the shape of the finding — causal, concentrated in the unbranded, moderated by volume — because two independent designs recovered it.

6.2 Healthcare-specific evidence

Xu, Armony and Ghose (2021) modelled review-derived service quality within a random-coefficient discrete-choice demand model applied to a leading United States appointment-booking platform, treating review content as a driver of physician demand rather than as a correlate of it. Text-derived quality proxies improved predictive accuracy of physician choice by 6–12% on mean squared error, both in and out of sample. Magnitudes beyond this sit behind the paywall; only the abstract was publicly retrievable.

Han, Lin, Han, Liao and Mei (2024), in the Journal of Medical Internet Research, report an experimental vignette study (n ≈ 461) on the effect of negative reviews and physician responses:

  • Exposure to detailed negative reviews reduced physician selection intention from 5.059 to 4.419 (p < 0.001).
  • A public physician response recovered most of the loss: 4.774 with a response present against 4.072 without (F₁,₄₅₇ = 21.849, p < 0.001), a recovery of approximately 0.70 against damage of approximately 0.64.
  • The review mix mattered considerably more than the reply. A 30% negative proportion produced selection intention of 3.760 against 5.053 at 10% negative (F₁,₄₅₇ = 83.618, p < 0.001). Standardised total effects were: negative-review proportion β = −0.371; claim type β = −0.343; physician response β = +0.194.
  • Specific factual complaints were substantially more damaging than vague ones: 3.805 for factual negatives against 5.020 for evaluative negatives (p < 0.001), a gap of 1.22 scale points. Concrete operational failures — waiting times, billing, process — cost more prospective patients than generic dissatisfaction.

The design measures stated intention in a hypothetical scenario rather than booked appointments, and should be read accordingly.

The operational reading is that public reply constitutes damage limitation worth approximately a third of what preventing or diluting the negative review would be worth, that the volume of genuine positive reviews is the primary lever, and that because factual complaints are the more damaging category, the operational failures generating them are where remediation pays back fastest.

6.3 Consumer thresholds and temporal decay

The following are drawn from BrightLocal (2026), a vendor-published consumer survey of general local businesses with no healthcare breakdown, conducted by a firm selling reputation software. They are directional and are not equivalent in evidentiary weight to the studies above.

  • Recency decays rapidly: 74% of respondents consider only reviews written in the preceding three months; 32% look to the preceding two weeks, up from 20% the prior year; 18% are influenced only by the preceding week.
  • Rating thresholds are rising: 31% report they will only use a business rated 4.5 or above, up from 17% within a single year; 68% will only use one rated 4.0 or above, up from 55%. A practice may therefore fall below the consideration threshold without any change in its own rating.
  • Review count operates as a filter rather than a gradient: 47% will not use a business with fewer than twenty reviews, and only 9% would use one with five or fewer.
  • Response behaviour matters: 80% report being likely to use a business that responds to all reviews; 42% report being unlikely to use one that never replies.
  • Solicitation is the mechanism: 78% were asked in the preceding twelve months; 83% of those asked left a review; 28% state they will "always" write one if asked, up from 16%. The source is internally inconsistent on this point, a chart caption on the same page giving 65% rather than 83%; the range 65–83% should be cited, or neither figure.

Whitespark (2026) reports an expert-panel elicitation in which forty-seven invited practitioners scored 187 ranking factors on a 0–5 scale. This is structured opinion rather than measurement, and no contributor has access to the ranking algorithm. Review attributes occupy five of the top nine conversion factors and both leading positions: high numerical rating (first), positive sentiment in review text (second), quantity of native reviews carrying text (fifth), and recency (seventh). "Sustained influx of reviews over time (rather than bursts)" is scored as a factor distinct from volume, constituting the practitioner case for continuous collection over campaign-based collection. No decay half-life or velocity threshold is quantified.

InsiderCX (2025), a vendor report drawn from a European customer base, reports a dose-response relationship between survey volume and Google review volume: clinics recording fewer than 500 survey responses average 95 Google reviews annually; 500–1,500 responses, 230; 1,500–3,000, 480; above 3,000, 790. This is correlational within a single vendor's customer base and patient volume is an evident uncontrolled confounder, since larger clinics both send more surveys and treat more patients. The pattern is consistent with the causal claim rather than evidence for it.

7. Service recovery

7.1 The service recovery paradox

Commercial literature commonly asserts that a recovered complainant becomes a more loyal advocate than a patient who never experienced a problem — the service recovery paradox.

De Matos, Henrique and Rossi (2007), in the Journal of Service Research, meta-analysed the paradox literature. The cumulative mean effect was significant and positive for satisfaction, but non-significant for repurchase intention, for word-of-mouth, and for corporate image. This directly falsifies the referral component of the commercial claim.

The meta-analysis was undertaken precisely because prior empirical tests disagreed, with only a subset supporting the paradox. Effects were moderated by study design, subject type and service category, and the categories examined comprised hotel, restaurant and "other." No healthcare or dental category was broken out, so the evidence base does not transfer to ambulatory practice without an explicit external-validity caveat. Numeric effect sizes sit behind the publisher's paywall and could not be verified; no specific magnitude should be cited.

7.2 The defensible ordering

AHRQ states a weaker proposition that does survive:

"The most satisfied customers are ones that have never experienced a serious problem… The next most satisfied are those who have experienced service difficulties — sometimes significant ones — that have been redressed by the organization. The least satisfied are those whose problems remain unsolved."

AHRQ does not claim that recovered patients exceed never-problem patients. The claim is one of ordering — resolved outperforms unresolved — and that is sufficient to justify the intervention. AHRQ's supporting "within a few percentage points" figure cites Goodman and Malech (1988) in The Quality Review, a trade publication, not peer-reviewed and not healthcare.

7.3 Asymmetry and the value of early interception

Richter and Muhlestein's finding that negative experience is more strongly associated with decreased profitability than positive experience is with increased profitability points in the same direction: the downside is the larger quantity. A programme optimised to identify detractors is better matched to the evidence than one optimised to raise a mean.

Combining the experimental results reported at 6.2: a factual negative review depresses selection intention by approximately 1.22 points more than a vague one, and a public reply recovers approximately a third of what preventing the negative would have been worth (β = +0.194 against −0.371). A complaint intercepted privately does not enter the denominator at all.

This is also the one component of the workflow that United States federal regulators expressly endorse. From the Federal Trade Commission's own guidance on the Consumer Reviews Rule: "Does the rule prohibit my company from contacting customers who post negative reviews to resolve the reported issues? No. It also does not prohibit simply asking satisfied customers to update their reviews."

8. The United States regulatory perimeter

The regulatory position runs opposite to the assumption most commonly held in practice. Systematic collection from all patients is expressly permitted and carries minimal exposure. The abbreviated alternative — screening patients and directing only those presumed satisfied to public review platforms — carries exposure on three independent axes.

8.1 HIPAA: collection is a permitted health care operation

45 CFR § 164.501 defines health care operations to include "(1) Conducting quality assessment and improvement activities, including outcomes evaluation…"

A covered practice may therefore use protected health information — patient identity, visit dates, contact details — to conduct patient-experience surveys concerning its own care without individual authorization. This constitutes a use of PHI rather than an exempt activity, but it is a permitted use.

Service recovery falls within the same permission. Section 164.501(6) enumerates "customer service" and "resolution of internal grievances" as health care operations; no separate basis is required to contact a dissatisfied patient. Section 164.501(5) covers use of aggregated data for business planning and cost-management analysis.

Two boundaries constrain this permission.

The research boundary. The operations permission holds only "provided that the obtaining of generalizable knowledge is not the primary purpose." Designing a programme as a systematic investigation intended to produce publishable or cross-practice generalizable findings converts it into research, requiring authorization or an IRB or Privacy Board waiver. Benchmarking or publication across practices is where this boundary is crossed.

The marketing boundary. A communication becomes marketing — requiring prior written authorization — where the practice receives financial remuneration from a third party whose product or service the communication describes. The trigger is third-party payment rather than promotional tone. A practice's own survey is not marketing; a survey mentioning a device or product brand that paid for the mention is. (§ 164.501 marketing definition, as amended by the 2013 HITECH Omnibus Rule.)

8.2 The Federal Trade Commission

A widely circulated claim requires correction here, and its correction matters because the overstatement invites the wrong remedy.

The FTC's 2024 Rule on the Use of Consumer Reviews and Testimonials (16 CFR Part 465, published at 89 FR 68077, effective 21 October 2024) does not contain a specific prohibition on review gating. From the Commission's own guidance: "Can my business ask for reviews only from customers whom we think are happy with our services? The rule does not contain a specific prohibition against such conduct. But this practice could violate the FTC Act. See, e.g., Endorsement Guides 16 C.F.R 255.2(d) and (e)(11)."

The exposure is real. It sits in Section 5 of the FTC Act and the Endorsement Guides, not in Part 465.

What Part 465 does provide is set out below.

Table 3. Substantive provisions of 16 CFR Part 465 and their application to ambulatory practice.

ProvisionProhibitionApplication to practice
§ 465.4Compensation or incentive conditioned, expressly or by implication, on a review expressing a particular sentimentA discount offered for a five-star review is squarely caught. The Commission's own examples of non-compliant wording mirror standard practice-marketing copy ("Tell us how much you loved your visit… and get a $5 coupon"). The wording of the solicitation, not merely its intent, is the compliance surface.
§ 465.7(a)Unfounded legal threats, physical threats, intimidation, or knowingly false public accusations used to suppress or remove a review; binds any person, not only businessesDirectly addresses the documented pattern of threatening a patient with defamation proceedings over an unfavourable review
§ 465.7(b)Materially misrepresenting that displayed reviews represent most or all reviews submitted, while suppressing by rating or sentimentNarrower than commonly reported. Carries an explicit safe harbour for withholding criteria applied equally to all reviews irrespective of sentiment (fake, off-topic, defamatory, containing another person's personal information)
§ 465.2(d)(1)The exemption on which the recommended design rests. Reviews resulting from generalized solicitations to purchasers are exempt. A practice soliciting every treated patient, conditioning nothing on sentiment, falls within it. A practice filtering solicitation by predicted sentiment forfeits the "generalized" characterisation.
§ 465.1(m)Defines "purchase a consumer review" to include non-cash consideration: gift certificates, products, services, discounts, coupons, contest entriesA treatment voucher or prize-draw entry offered for a review constitutes purchasing a review
§ 465.1(d)Defines "consumer review" to include bare star ratings carrying no textA star-only solicitation falls inside the Rule, not outside it

Enforcement is not theoretical. In January 2022 the Commission's first action concerning concealment of negative reviews resulted in Fashion Nova paying $4.2 million and being prohibited from suppressing reviews. The charged mechanism is precisely the gating pattern: automatic publication of four- and five-star reviews while lower-rated reviews were held for an approval that did not occur, with hundreds of thousands of lower-starred reviews unpublished between late 2015 and November 2019.

Three details warrant note. Liability extends to the vendor layer: the Commission simultaneously notified ten review-management service companies that avoiding the collection or publication of negative reviews violates the FTC Act, so a practice's software provider's design is the practice's exposure. The permissible grounds for withholding are content-based only — obscene, sexually explicit, racist, unlawful, or unrelated — and rating level is not among them. On penalties, the Commission placed over 700 businesses on notice of civil-penalty exposure in October 2021, each violation of a final Commission order then carrying up to $46,517; Part 465 authorises courts to impose civil penalties for knowing violations, but the Commission's guidance publishes no per-violation figure and identifies itself as non-binding staff guidance conferring no safe harbour.

8.3 The Telephone Consumer Protection Act

47 CFR § 64.1200 governs automated survey contact. Three findings are material, one of which reverses a common assumption.

A pure survey is not telemarketing, so the consent threshold is lower than generally assumed. Prior express written consent is required only for autodialed or prerecorded calls and texts that "include or introduce an advertisement or constitute telemarketing" (§ 64.1200(a)(2)). A patient-experience survey carrying no marketing content falls under (a)(1), requiring prior express consent — a materially lower threshold than signed authorization.

The healthcare exemption does not cover surveys. The FCC's health-care-provider exemption at § 64.1200(a)(9)(iv)(C) is a closed list of nine permitted purposes: appointment and examination confirmations and reminders, wellness checkups, hospital pre-registration instructions, pre-operative instructions, laboratory results, post-discharge follow-up to prevent readmission, prescription notifications, and home healthcare instructions. Satisfaction surveys, recommendation-intention requests and feedback solicitation do not appear on it. A practice cannot rely on the healthcare exemption to send a survey without consent. Subsection (a)(9)(iv)(D) additionally bars solicitation or advertising content and requires HIPAA compliance.

Bundling an offer reclassifies the message. Section 64.1200(f)(13) defines telemarketing by purpose: any call or message initiated to encourage a purchase. A feedback survey carrying a rebooking offer, treatment upsell or membership proposition becomes telemarketing and triggers the signed written-consent requirement.

Two further requirements apply. Section 64.1200(a)(10) requires revocation to be honoured within a reasonable period not exceeding ten business days, requires any reasonable method of revocation to be accepted (with "stop," "quit," "end," "revoke," "opt out," "cancel," and "unsubscribe" per se reasonable), and expressly forbids designating an exclusive opt-out channel; instructions confining revocation to a single method are non-compliant. And § 64.1200(a)(3)(v) imposes a hard frequency cap on prerecorded voice messages to residential lines: HIPAA-related health care messages are exempt from consent only within one call per day and a maximum of three per week per patient, and only where opt-outs are honoured.

8.4 Platform policy

Review gating is a platform violation independently of federal regulation. Google Maps' user-contributed-content policy prohibits, in its own terms, conduct that would "discourage or prohibit negative reviews, or selectively solicit positive reviews from customers."

It separately prohibits incentives of any kind — "payment, discounts, free goods and/or services" — in exchange for posting, revising or removing a review; pressuring patients to review while on the premises, which excludes the on-site tablet pattern; and requesting that specific content be included, which excludes the widespread practice of asking patients to name a particular clinician.

What the policy affirmatively permits is the design this review recommends: to "solicit or encourage the posting of content that does represent a genuine experience, without offering incentives to do so or attempting to influence the rating or the contents of the review."

Enforcement escalates beyond the individual review: the policy states that Google "may take actions that will range from suspending the account privileges to account termination." For most practices the Google Business Profile is the single largest local-acquisition asset. Whitespark's (2026) expert panel places reports of review gating sixth in its scored suspension-risk ranking, above reports of fake reviews.

8.5 An unresolved question

This review sought primary sources on state recording-consent requirements applicable to telephone surveys and did not locate them. The general position — that a substantial minority of states require all-party consent to record a call, and that a practice recording survey or recovery calls should disclose and obtain consent to the strictest applicable standard, typically that of the recipient's jurisdiction — is widely stated but is not sourced here, and should be confirmed with counsel before any call recording is enabled.

9. Operational parameters

9.1 Timing

Johnston et al. (2021), across eighty-seven primary care practices, triggered automated surveys within seventy-two hours of the visit and achieved 55.6% response among consenting patients.

A vendor source (rater8, 2026) asserts a twenty-four-hour window and states that surveys sent at forty-eight hours "lose the majority of potential respondents." No data, citation or decay curve is offered for either figure. A same-day to seventy-two-hour window is defensible on the peer-reviewed evidence; the specific twenty-four-hour cliff is not.

9.2 Mode, and the confound within it

Johnston et al. report, in a pre-consented waiting-room-recruited cohort: email 60.9% (369/606) against automated telephone 38.1% (101/265), p < 0.001, holding regardless of age, sex or chronic-disease status. The authors' own literature review places the realistic benchmark for cold emailed surveys in primary care at 20–30%; the 60.9% figure reflects pre-consent and should not be used for planning.

The more consequential finding, and the one least frequently accounted for, is that survey mode alters the responses. The same patients answering the same items scored differently by channel. Against a paper waiting-room baseline of 1.26 on an item concerning the provider's access to test results, telephone respondents averaged 1.97 and email respondents 1.65 (both p < 0.001), with telephone respondents "particularly critical with respect to care coordination."

The consequence is that scores collected through different channels cannot be pooled or trended against one another. A practice migrating from email to SMS and observing a score movement has learned nothing about its care. Channel should be changed once, deliberately, with the trend line restarted.

A vendor quasi-experiment (rater8, 2026) reports 164 treatment practices switching from email-only to text-only against 531 controls over a ninety-day pre-post window (March–June 2025): treatment response rose from 22.40% to 37.52%, a 67% relative increase, while the already-texting control moved from 34.74% to 35.59% — a channel-attributable lift of approximately fifteen percentage points. No significance test, confidence interval or specialty breakdown is reported, the study is non-randomised, and the publisher sells the intervention. It is marketing-grade evidence.

The same source reports that the channel change did not degrade response quality: comment rate rose from 38.03% to 41.74%, mean comment length held at approximately 81–84 characters, and mean star rating moved from 4.90 to 4.92. Subject to the same evidentiary caveats, this is a useful counter to the concern that surveying a larger proportion of patients will depress the average.

InsiderCX (2025), drawn from a European sample, reports for SMS and WhatsApp post-visit surveys: 90.2% delivery (310,047 of 343,764 invitations), a 21.2% start rate on delivered invitations, and 83.3% completion among those started, with start rates by specialty ranging from 24.9% (fertility) to 16.4% (aesthetics). The report's own cited United States comparator of 23% is higher than its result despite favourable presentation, so 21.2% should be treated as a floor rather than as evidence of channel superiority. The source carries an unexplained internal discrepancy between 52,540 and 54,757 responses.

Measured levers from the same vendor's A/B testing: a reminder raises effective response from 22.5% to 28.4% (+5.9 percentage points); send-time optimisation from 20.6% to 23.5% (+2.9 points); a 13:00 send slot achieves 24.6% against 18–20% for most other hours.

9.3 Instrument length

Johnston et al. report that a five-question survey produced 97.1% completion among those who started (470/484). Abandonment is not the binding constraint at five items; initiation is.

9.4 Published benchmarks

InsiderCX (2025) reports a global healthcare NPS of 79.6 (±0.86), with dental at 83.8, aesthetics at 76.5, fertility at 88.6, diagnostics at 83.8, polyclinics at 80.3, and surgical at 66.8. The sample is European and drawn from the vendor's own customer base.

Read against the vendor's evident intent, this cuts the other way. If substantially every practice scores in the high seventies and eighties, the score has little power to discriminate between practices. Adams et al. (2022) reach the same conclusion from the peer-reviewed data, reporting NPS varying from 44 to 83 across different hand surgeries alone, and 71 for total hip against 49 for total knee replacement, and concluding that NPS suits localised performance assessment rather than benchmarking between organisations.

Published specialty benchmark figures should therefore be treated as a marketing artefact of which patients were surveyed, in which country, concerning which procedures. A practice's own prior-period score is the only comparator with defensible validity.

A further vendor figure requires qualification: rater8 cites a national average patient-experience survey response rate of approximately 23%, attributed to the HCAHPS Toolkit (Flex Monitoring Team, August 2025). HCAHPS is a CMS inpatient instrument and is not a like-for-like norm for outpatient private practice.

9.5 Non-response bias

Johnston et al. found respondents skewed older, more female, higher-income and higher-morbidity. Email response rose with income (57.4% to 66.3% across the upper four quintiles) while telephone response did not (36% to 43%), the two converging only in the lowest quintile (both 46%).

A score therefore describes respondents rather than patients. The response rate should be reported alongside any score. A score reported without its denominator is not interpretable.

10. Discussion

10.1 Synthesis

The argument for systematic collection is stronger when it is narrower. Three findings, none contingent on the validity of any particular metric, converge.

First, the count of patients reporting a problem is a validated risk signal (Hickson et al., 2002), and it is unobservable without a channel that invites it (AHRQ). Second, displayed ratings causally affect revenue, and do so most for the smallest, least-branded and lowest-volume businesses — the population most likely to conclude that it is too small to warrant the effort (Luca, 2016; Anderson and Magruder, 2012). Third, reputation decays within weeks and is driven predominantly by solicitation (BrightLocal, 2026; Han et al., 2024), so the operative choice is not between a feedback process and none, but between a process the practice conducts and one conducted by whichever patients were sufficiently dissatisfied to post unprompted.

The regulatory position reinforces rather than complicates this. HIPAA expressly permits collection, follow-up contact, and internal analysis. What carries risk — federal penalty, platform termination, and a $4.2 million enforcement precedent — is the abbreviated alternative of soliciting only patients expected to respond favourably.

The metric question, which dominates commercial discussion, is largely immaterial. Keiningham et al. found that replacing the single recommendation item with multi-item loyalty measures increased explanatory power by only 0.1 to 1.6 percentage points of R²; East et al. (2011) found NPS and the ACSI correlating at r = 0.99, 0.72 and 0.92. The critique establishes that NPS is not the best available predictor, not that it is materially worse than the alternatives. Whether a structured signal is collected at all matters considerably more than which instrument collects it.

The defensible position is therefore that the eleven-point item is an adequate mechanism for triggering an action and a poor mechanism for measuring a practice: a router rather than a key performance indicator, with free text occupying the position the score conventionally holds. This is not merely a rhetorical resolution. Xu et al. (2021) found that for a given service feature, the free-text description is a better measure of perceived quality than the numeric rating of it, and Adams et al. independently identify the comments section as "the most useful component of NPS surveying," beneficial in four of twelve reviewed studies.

One qualification attaches to the last point. Adams et al. state that comments are "typically completed by more than three-quarters of respondents." This figure should not be used. It carries exactly one citation — a United Kingdom pilot in three primary-care oral surgery practices (Gerrard, Jones and Hierons, 2017) — and the wider literature is considerably lower: the English GP Patient Survey recorded a 44.4% comment rate (3,426 of 7,721 patients across twenty-five practices) and the Swiss SCAPE survey 31% (844 of 2,755). Planning should assume approximately one third to one half of respondents leaving a comment.

10.2 A minimum specification

The evidence supports a specification of the following minimum form. Each element is included because a finding above requires it.

Instrument and delivery. A single message to every patient between twenty-four and seventy-two hours after the visit, with no sentiment filter at any point. Universal solicitation is what places the practice within the FTC's generalized-solicitation exemption at § 465.2(d)(1) and within Google's permitted-solicitation language, and it is the only design the Commission's own guidance endorses. Two items: an eleven-point rating followed by a free-text prompt. The free text is the diagnostic payload; the number is a routing trigger. No offer, upsell or promotional content, since bundling one reclassifies the message as telemarketing under 47 CFR § 64.1200(f)(13) and raises the consent threshold to signed written consent.

Routing. Any score of 0–6, and any negative free-text response at any score, should alert a single named individual the same day. The second condition is not redundant: passives (7–8) are excluded from the NPS calculation entirely, yet InsiderCX report that 37.8% of respondents scoring 7–8 left negative or mixed written feedback against 6.7% of promoters, a 5.6-fold difference. Routing should follow the text, not the band. Service recovery is an enumerated health care operation under 45 CFR § 164.501(6), so no separate authorization is required for the contact.

Review invitation. The same invitation should be sent to every respondent, in the same message, irrespective of score. This single decision maintains compliance on all three axes — Section 5 of the FTC Act, Part 465, and platform policy — and the evidence indicates that solicitation is what drives review volume in any case.

Measurement. Three figures monthly: the count of detractors, not the mean, since the count carried the predictive signal in Hickson et al.; the response rate, reported alongside anything else; and three to five recurring free-text themes, with waiting time, billing clarity and process failures prioritised because Han et al. found factual complaints more damaging than evaluative ones by 1.22 scale points. A score, if reported at all, should be reported as a rolling twelve-month figure with its credible interval, and not at all below approximately 160 responses per window (Costa and Ponte, 2024).

Consent. A single TCPA consent provision on the intake form covering non-marketing SMS from the practice. A pure survey requires prior express consent rather than signed written consent under § 64.1200(a)(1), but the healthcare exemption at (a)(9)(iv)(C) is a closed list that does not include surveys, so consent is required; revocation must be honoured within ten business days by any reasonable method.

Contraindicated practices. Soliciting only patients presumed satisfied (FTC Act § 5 exposure via Endorsement Guides 16 CFR § 255.2(d), (e)(11); platform violation escalating to profile termination; $4.2 million precedent). Offering anything of value for a review, including discounts, vouchers, products or prize-draw entries (16 CFR § 465.1(m) covers non-cash consideration; platform policy prohibits it outright). Wording a solicitation so that a favourable review is implied. Presenting a review interface to a patient on the premises, or requesting that a named clinician be mentioned. Comparing a practice score to a published specialty benchmark, or to the practice's own score obtained through a different channel.

10.3 Limitations

This review is narrative rather than systematic, and its search strands were constructed to answer a practitioner-facing question rather than to achieve exhaustive coverage. No formal risk-of-bias instrument was applied.

Several load-bearing sources could not be read in full. Richter and Muhlestein (2017), Xu et al. (2021) and De Matos et al. (2007) were available only as abstracts, and no magnitudes are reported from them.

The strongest causal evidence in the review — Luca (2016) and Anderson and Magruder (2012) — is drawn from restaurants, not healthcare. The transfer of the qualitative shape is defensible on the strength of two independent replications; the transfer of coefficients is not, and no coefficient is transferred here.

The healthcare NPS literature is predominantly British and publicly funded, which limits its applicability to United States private practice on both reimbursement and consumer-behaviour grounds.

Vendor-published material is included because it constitutes the evidence practitioners encounter, but it is drawn from single-vendor customer bases, is frequently non-randomised, and in two instances contains unexplained internal inconsistencies noted at the point of citation.

Finally, the regulatory analysis states the position as at 3 September 2026 and is not legal advice. Section 8.5 identifies a question this review could not resolve from primary sources.

11. Conclusion

The proposition that a practice should collect patient feedback because a recommendation-intention score predicts growth is not supported by the literature, and advancing it to a sceptical audience forfeits credibility that the substantive argument does not require.

The substantive argument is that the count of dissatisfied patients is a validated predictor of risk and is invisible without a channel that solicits it; that displayed ratings causally move revenue, and do so most for precisely those practices that judge themselves too small to be affected; that reputational standing decays within weeks and responds principally to being asked; and that the regulatory framework permits universal collection while penalising selective solicitation.

The metric is not the intervention. The mechanism is: ask every patient, keep the instrument to two items, treat the free text as the finding and the number as a trigger, count detractors rather than trending a mean, and act within the day.

References

  1. Adams, C., Walpola, R., Schembri, A.M. and Harrison, R. (2022) 'The ultimate question? Evaluating the use of Net Promoter Score in healthcare: A systematic review', Health Expectations, 25(5), pp. 2328–2339. PMC9615049. Peer-reviewed.
  2. Agency for Healthcare Research and Quality (2022) CAHPS Improvement Guide, Strategy 6P: Service Recovery Programs. Content last reviewed April 2022. Government guidance.
  3. Anderson, M. and Magruder, J. (2012) 'Learning from the Crowd: Regression Discontinuity Estimates of the Effects of an Online Review Database', The Economic Journal, 122(563), pp. 957–989. Peer-reviewed.
  4. Baehre, S., O'Dwyer, M., O'Malley, L. and Lee, N. (2022) 'The use of Net Promoter Score (NPS) to predict sales growth: insights from an empirical investigation', Journal of the Academy of Marketing Science, 50(1), pp. 67–84. Peer-reviewed.
  5. BrightLocal (2026) Local Consumer Review Survey 2026. Vendor-published; general local-business sample; no healthcare breakdown; internally inconsistent on one reported figure.
  6. Costa, E.G. and Ponte, R.T.Q. (2024) 'Minimum sample size for estimating the Net Promoter Score under a Bayesian approach', Anais da Academia Brasileira de Ciências, 96(2), e20230991. Peer-reviewed.
  7. De Matos, C.A., Henrique, J.L. and Rossi, C.A.V. (2007) 'Service Recovery Paradox: A Meta-Analysis', Journal of Service Research, 10(1), pp. 60–77. Peer-reviewed; abstract only.
  8. Doyle, C., Lennox, L. and Bell, D. (2013) 'A systematic review of evidence on the links between patient experience and clinical safety and effectiveness', BMJ Open, 3, e001570. PMC3549241. Peer-reviewed.
  9. East, R., Romaniuk, J. and Lomax, W. (2011) 'The NPS and the ACSI: A critique and an alternative metric', International Journal of Market Research, 53(3). Peer-reviewed; cited via Sauro (2019).
  10. Federal Trade Commission (2022) Fashion Nova Will Pay $4.2 Million as Part of Settlement, press release, 25 January 2022. Regulatory — enforcement record.
  11. Federal Trade Commission (2024) Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465. 89 FR 68077, 22 August 2024; effective 21 October 2024. Regulatory.
  12. Federal Trade Commission (2024) Consumer Reviews and Testimonials Rule: Questions and Answers, November 2024. Non-binding staff guidance; confers no safe harbour.
  13. Fisher, N.I. and Kordupleski, R.E. (2019) 'Good and bad market research: A critical review of Net Promoter Score', Applied Stochastic Models in Business and Industry, 35, pp. 138–151. Peer-reviewed; Practitioner's Corner — critical review presenting no new data; second author has a competing-methodology interest not declared as a limitation.
  14. Gerrard, G., Jones, R. and Hierons, R.J. (2017) 'How did we do? An investigation into the suitability of patient questionnaires in three primary care oral surgery practices', British Dental Journal, 223(1), pp. 27–32. Peer-reviewed.
  15. Goodman, J. and Malech, A. (1988) The Quality Review. Trade publication; not peer-reviewed; not healthcare. Cited via AHRQ.
  16. Google (2026) Maps user-contributed content policy: prohibited and restricted content. Retrieved 29 August 2026. Regulatory — platform policy.
  17. Han, X., Lin, Y., Han, W., Liao, K. and Mei, K. (2024) 'Effect of Negative Online Reviews and Physician Responses on Health Consumers' Choice: Experimental Study', Journal of Medical Internet Research, 12 March 2024. PMC10966444. Peer-reviewed.
  18. Hickson, G.B., Federspiel, C.F., Pichert, J.W., Miller, C.S., Gauld-Jaeger, J. and Bost, P. (2002) 'Patient Complaints and Malpractice Risk', JAMA, 287(22), pp. 2951–2957. Peer-reviewed.
  19. InsiderCX (2025) 2025 Benchmark Report, Volume 01. Vendor-published; European sample drawn from vendor customer base; two unexplained internal discrepancies.
  20. Johnston, S., Hogg, W., Wong, S.T., Burge, F. and Peterson, S. (2021) 'Differences in Mode Preferences, Response Rates, and Mode Effect Between Automated Email and Phone Survey Systems for Patients of Primary Care Practices', Journal of Medical Internet Research, 11 January 2021. PMC7834947. Peer-reviewed.
  21. Keiningham, T.L., Cooil, B., Andreassen, T.W. and Aksoy, L. (2007) 'A Longitudinal Examination of Net Promoter and Firm Revenue Growth', Journal of Marketing, 71(3), pp. 39–51. Peer-reviewed.
  22. Luca, M. (2016) Reviews, Reputation, and Revenue: The Case of Yelp.com. Harvard Business School Working Paper 12-016, originally 2011, revised March 2016. Peer-reviewed working paper.
  23. Mecredy, P., Wright, M.J. and Feetham, P. (2018) 'Are promoters valuable customers? An application of the net promoter scale to predict future customer spend', Australasian Marketing Journal, 26(1). Peer-reviewed; cited via Sauro (2019), primary source not retrieved.
  24. Pingitore, G., Morgan, N.A., Rego, L.L., Gigliotti, A. and Meyers, J. (2007) 'The single-question trap', Marketing Research, 19(2). Peer-reviewed; cited via Sauro (2019).
  25. rater8 (2026) Patient survey response rate article, May 2026. Vendor-published; non-randomised quasi-experiment; no significance testing reported.
  26. Reichheld, F.F. (2003) 'The One Number You Need to Grow', Harvard Business Review, December 2003. Practitioner publication.
  27. Richter, J.P. and Muhlestein, D.B. (2017) 'Patient experience and hospital profitability: Is there a link?', Health Care Management Review, 42(3), pp. 247–257. Peer-reviewed; abstract only, closed access.
  28. Rocks, B. (2016) 'Interval Estimation for the "Net Promoter Score"', arXiv:1601.07235. Preprint; results reproduced in peer-reviewed literature.
  29. Sauro, J. (2019) Has the Net Promoter Score Been Discredited? MeasuringU, October 2019. Secondary review; source for East et al. (2011), Pingitore et al. (2007) and Mecredy et al. (2018).
  30. Stimson, C.J. et al. (2010) 'Medical malpractice claims risk in urology: an empirical analysis of patient complaint data', Journal of Urology. Peer-reviewed; cited via verification, not read in full.
  31. United States. 45 CFR § 164.501 — HIPAA definitions: health care operations, marketing, research. 65 FR 82802 as amended through 78 FR 5695 (2013 HITECH Omnibus Rule). Regulatory.
  32. United States. 47 CFR § 64.1200 — TCPA implementing rules, current through 90 FR 42138. Regulatory.
  33. United States. 16 CFR § 255.2(d), (e)(11) — Endorsement Guides. Regulatory.
  34. Whitespark (2026) Local Search Ranking Factors 2026. Elicitation of 47 practitioners scoring 187 factors. Expert panel; explicitly not measurement.
  35. Xu, Y., Armony, M. and Ghose, A. (2021) 'The Interplay Between Online Reviews and Physician Demand: An Empirical Investigation', Management Science, 67(12), pp. 7344–7361. Peer-reviewed; abstract only, closed access.

Appendix A — Claims examined and not supported

The following claims circulate widely in commercial literature on this subject. Each was examined and is not asserted in this review. They are listed so that the boundary of the evidence is visible rather than implicit.

Table 4. Claims examined during this review and not asserted, with the reason for exclusion.

ClaimReason it is not asserted
NPS predicts practice growthNot detected by simple correlation in two independent replications (Keiningham et al., 2007; Baehre et al., 2022)
Recovered complainants become the strongest referrersMeta-analytically non-significant for word-of-mouth and repurchase intention (De Matos et al., 2007)
Complaint content identifies the high-risk clinicianHickson et al. (2002) report that "the total number of patient complaints, not any particular type" predicted risk-management outcomes
96% of dissatisfied patients tell nine or ten othersUncited on the AHRQ page and unsupported by any of that page's five numbered references
A specialty NPS threshold constitutes a targetCross-organisation benchmarking is unsupported; scores vary from 44 to 83 by procedure alone (Adams et al., 2022)
Three-quarters of respondents leave a commentTraces to a single UK pilot in three practices; the wider literature reports 31–44%
Surveys must be sent within twenty-four hoursAsserted by a vendor with no data, citation or decay curve
A one-star gain is worth 5–9% to a given practiceThe causal estimate is from restaurants; the shape transfers, the coefficient does not
Effect sizes for treatment-plan acceptance or non-attendanceNot located in the peer-reviewed literature by this search
Specific state call-recording consent requirementsNo primary source surfaced; confirm with counsel before enabling recording

Appendix B — Verification summary

Research conducted 29 August 2026 via a five-strand parallel search. Twenty-seven sources were retrieved and read. One hundred and thirty-five candidate claims were extracted, of which twenty-five were subjected to three-vote adversarial verification. Eighteen claims were confirmed at 3–0 or 2–1; seven were refuted and are either excluded or reproduced in Appendix A with the reason for exclusion. Where a source could not be read in full — paywalled abstracts and closed-access bodies — that limitation is stated at the point of citation and no magnitude is reported from that source.

Want this kind of thinking applied to your practice?

Twenty minutes with us. We'll audit your current review velocity and tell you honestly whether applaud fits.