All posts
Writing12 min read

Making the score move the practice, not the other way round

Most practices that adopt NPS measure it well and act on it never. The evidence says the number adds little on its own — so stop managing to it. Use it to sort the week's conversations, then run the themes: one theme, one owner, one week.

Tom North·Founder, applaud·

A practice adopts NPS in the spring. By early summer the survey is running properly: it goes out after every visit, the response rate is respectable, the sample is not just the four people who wanted to complain. There is a number, and it appears in the Friday email with an arrow beside it. Everyone nods at the arrow. Eighteen months later the number is within a point of where it started, and nobody in the building can name one thing the practice does differently because of it.

That is not a measurement failure. The measuring worked fine. It is an ownership failure, and the two look identical from the outside, which is why the usual response is to go shopping for a better survey.

The number, on its own, does not move anything

The question comes from Fred Reichheld, writing in Harvard Business Review in December 2003. He proposed “would you recommend this company to a friend?” as the best single-question predictor of top-line growth, on the back of two years of research tying survey answers to actual purchasing and referral behaviour.[1] Bain, which runs the system, states the mechanics precisely: the question is “how likely is it that you would recommend [product, service or company/brand] to a friend or colleague?”, answers of 9–10 are promoters, 7–8 are passives, 0–6 are detractors, and the score is the percentage of promoters minus the percentage of detractors.[2]

Then healthcare got hold of it. The most useful summary of what happened next is a 2022 systematic review in Health Expectations by Adams and colleagues, which found 12 studies meeting inclusion criteria.[3] Four of them highlighted advantages: ease of use, high completion rates, and being well understood by a range of patients. Three raised concerns about whether the recommendation question is even relevant, particularly when patients cannot choose their provider, which is the inpatient and single-payer setting the review is mostly describing. And four of the twelve determined that NPS adds minimal value to healthcare improvement. The review concluded that NPS should be used in conjunction with a larger survey rather than standalone.

Krol and colleagues had reached a related conclusion in the same journal in 2015: patient experiences correlated more weakly with the NPS than with a plain global rating.[4] I went through that literature properly in an earlier post on what the score is worth. The short version: it is weaker than its reputation, and least weak in the settings where patients pick their own provider, which is most of outpatient medicine.

And the growth claim, the one that made NPS famous, has not held up under replication. Keiningham and colleagues could not reproduce the claimed superiority of Net Promoter over other measures, in the very industries that had been cited as its exemplars.[5] So let me put this bluntly, because I have heard it said to practice owners in rooms I was in: nobody has shown that your practice’s NPS predicts your revenue. The one serious replication attempt could not reproduce the claim in the industries the original research picked as its best cases, and no study I can find has tested it in outpatient medicine at all. Anyone who tells you the number forecasts your growth is ahead of the evidence.

We sell a product with an NPS score in it. That body of research is the strongest available argument against half of what we sell, and I would rather you read it than not. But read what the reviews actually concluded, which is not “stop measuring.” It is that the number is weak as a standalone verdict on a practice. That is a statement about how you use it.

What the number is actually for

Look at the formula again: promoters minus detractors, as a percentage. The passives are not in it. They are collected, counted into the denominator, and then cancelled out by construction. A single figure that discards its own middle band, aggregated over a week of conversations that were each about something specific, is not a diagnosis. It is a sorting key.

Every week a practice receives a pile of patient conversations. The band a patient lands in tells you what to do with that one conversation, and nothing more:

  • 0–6: somebody calls them. A person, from the practice’s own number, inside a day, who is allowed to fix something.
  • 7–8: somebody reads what they wrote. No call, no form, no follow-up survey. Read it.
  • 9–10: a review link, and nothing else asked of them.

Once the number has routed the conversation, it has done everything it is good at. Report it if the owners like seeing it, but do not set a target on it and never tie a bonus to it, because the fastest way to move it is to change who you survey.

One thing here matters more than it sounds: the band decides where the conversation lands, never who gets asked for a review. Nobody is screened out of the invitation because of the answer they gave. That distinction is the whole subject of a separate post, and it is the first thing I would check in any software you are being sold. The one exception is ordinary courtesy — a patient in the middle of a complaint does not get an automated review ask on top of it.

The themes are the product

I pointed at this in the earlier post and moved on. It deserves more than a mention, because it is the only part of the instrument this post asks you to do anything with. The Adams review says it outright: the most useful component of NPS surveying was the patient comments section, viewed positively in four of the 12 studies.[3] The part of the instrument people actually valued is the part that is not a number. The same review reports the objection, and it is a fair one: clinicians in some of those studies found the comments too thin to act on, short on the cause of the dissatisfaction and short on what to do instead. That is a design problem rather than a reason to stop reading them, and it is why the vocabulary is fixed and the quotes are verbatim.

So make the sentences the working artefact. applaud’s platform reads each answer for sentiment, intent, risk and theme, and the themes come from a fixed twelve-term vocabulary — the front desk, the wait, being kept informed, billing clarity, and so on. Fixed matters more than it looks. A theme list that grows every week cannot be counted week over week, and if “waiting”, “wait time” and “kept waiting” are three separate rows, the largest problem in the practice is spread thinly enough to look like three small ones.

The other constraint is verbatim. A quote is kept only if it appears character-for-character in what the patient actually wrote. The reason is what happens in the room. “Patients expressed frustration around billing communication” invites the reply “which patients, and frustrated how?”, and ten minutes go on whether the summary is fair. A sentence a patient typed with their own thumbs ends that argument before it starts. It also constrains the model: a paraphrase drifts toward whatever the summariser expected to find, and nobody in the meeting can tell. Character-for-character is a property somebody can check.

The practical output is three theme names, a count under each, and four or five sentences a patient wrote. In our case it lands behind the practice’s login with a report every Friday, and the answers come from a conversation Robin had by text or phone rather than a form nobody opens.

The operating rhythm: one theme, one owner, one week

This is the part worth stealing whether or not you ever buy anything from us. Five rules, and they only work together.

  1. The report shows the top three themes by volume. Three. Not ten. A ten-row table gets read to row four and skimmed after that, and the rows nobody reads are not selected at random — they are the ones at the bottom, which is where a rising problem sits before it is a big one. If you do not have software grouping the comments, this is the only rule that costs you anything: one person reads the week’s free text and tallies it against a list of theme names you write once and do not change. Twenty minutes, and it is the twenty minutes that makes the other four rules possible.
  2. Each of the three gets a named owner before the meeting ends. Not a department. A name. Somebody who was in the room, heard their name, and knows it is theirs until next Friday.
  3. A theme nobody will own is struck, and stays out for four weeks. Not deferred, not carried in grey — out, even if its volume would put it back in the top three. If it is still there in week five it comes back, and somebody takes it or the practice writes down why not.
  4. The owner’s job is one change, testable, described in a sentence. Not a project, not an initiative, not a policy review. If it needs a plan, it is not this week’s change; break off the part that does not.
  5. Four weeks on, the theme is either lower than it was or it is not. A single week’s movement is not evidence and nobody at the table should argue from it. If a theme is running under about ten comments a week, do not read week-to-week movement at all.

Rule three is the one people argue with, so let me defend it. Carrying an unowned row costs nothing in the moment, which is precisely why it is the most expensive habit in the rhythm. A theme that has appeared in eleven consecutive reports with nobody’s name against it has stopped being information and become wallpaper. Worse, it teaches its readers something true: this document contains rows nobody intends to act on. The rational response to that is to skim. And once a reader skims a report, they skim all of it, including the row that mattered this week.

Deleting the unowned row does not mean giving up on the problem. It writes down a decision the practice had already made in silence — that this is not what we are working on — and puts it somewhere a partner can look at it and disagree. Half the value of the rule is the number of times somebody says “wait, do not strike that one, I will take it” in the ten seconds before it disappears.

On rule five, one honest caveat. A single week of movement on a small comment count is noise, and if you treat it as proof you will chase your own sampling error around the calendar. What you are reading is direction across four or five weeks. The more reliable signal is the absence of movement across eight, and I will come back to that.

Here is what the sheet looks like when it is doing its job.

ThemeShare of commentsOwnerThe one change this week
The wait31%Practice managerAnyone whose provider is running more than 20 minutes behind gets a text before they leave home.
Billing clarity18%Billing coordinatorThe out-of-pocket figure is said out loud at check-in and written on the visit summary.
Being kept informed14%Lead assistantWhoever takes a patient back states the current run-time and the reason, in one sentence.

Table 1. An illustrative weekly theme sheet. These are not figures from a real practice; they are made up to show the shape. In a real sheet the owner column holds first names rather than job titles, and a fourth theme that appeared last week without an owner is simply absent from this one.

The meeting that produces it runs fifteen minutes, because fifteen minutes is what survives contact with a Friday afternoon. Two minutes reading the three theme names and their quotes aloud, in the patient’s own words. Three minutes assigning owners. Five between them for each owner to say their one change in a sentence while somebody writes it in the cell. Five on the running four-week direction of the three: is any of them lower than when its owner took it, and if not, what is the second thing to try. Then it is over.

Read the passives

Now a rule I hold strongly and cannot cite. This is our opinion from watching theme distributions, not a finding from the literature, and nothing I have read tests it. If it turns out to be wrong I would want to know.

Read the 7s and the 8s. When the theme sheet gives you three rows, read the passive comments under each row first. That band has the highest ratio of fixable operational friction to unfixable clinical or personal circumstance in the whole survey. The parking. The twenty minutes in the waiting room with nobody explaining why. The bill that was not the number they had in their head. None of it individually is worth making a complaint about, which is exactly why it surfaces at a 7 and not at a 3 — and it is also the category a practice can change inside a week without spending money.

Compare that to the other two bands. Detractors need a phone call, and that call is service recovery, a different job with a different skill. Much of what produces a detractor is not front-desk friction at all: an outcome, a diagnosis, a cost an insurer decided, a six-week wait that is a capacity problem rather than a process one. Worth the call. Rarely worth a Tuesday change. Promoters need a review link and nothing else asked of them, and the reason that matters is arithmetic rather than sentiment — in our own 2026 dataset of 15,061,473 Google reviews, the median US practice holds 19 reviews in total.

Which leaves the passives with nobody looking at them. There is a structural reason for that, and it is the formula: promoters minus detractors leaves the 7s and 8s out of the score entirely.[2] The band with the most actionable content is the band the metric is built to discard. I do not read that as an argument for changing the formula. I read it as one more reason the number is a routing device and the comments are the finding.

The detractor who never replies

Some will not talk to you. Set the rule in advance so nobody has to decide in the moment: one call attempt, from the practice’s own number, within a day, made by somebody with the authority to actually resolve something. A voicemail that names a person and gives a direct line. Then stop. Chasing a patient who has decided they are done converts a bad experience into a worse one, and you will read about it later in public.

Then log the non-reply as a fact of its own. If most of a practice’s detractors never pick up, that is a finding about the call rather than about the patients — the number it comes from, the hour it goes out, whether the voicemail says a name. A theme, even though no patient wrote it. And unless a complaint is open they still receive the ordinary invitation on the normal schedule: day two and day five reminders, then nothing, and the sequence stops the moment the link is clicked or a review lands. Withholding the ask because somebody scored low is the thing we refuse to do, and the reason is Google’s own policy on selectively soliciting positive reviews — which I go through properly in the post on routing and gating.

Eight weeks, no movement

Eight weeks in, one row will not have budged, and there are two honest explanations. Either the change was wrong, or the theme is not fixable at the level you are trying to fix it. The second is more common than anyone wants to admit. Wait times in a practice deliberately booked to a revenue target are not a front-desk problem; they are a schedule design decision, and no quantity of apologising at the window will move the comment count. Billing confusion in a practice with four plan types and a two-page estimate is not a coordinator problem either.

So put a stop rule on it. After four weeks and two genuinely different changes by the same owner, one of two things happens: the constraint goes up a level, to whoever controls the schedule or the software or the staffing, or the theme is retired with a written sentence — “we have chosen to run this way and this is the cost.” Retiring it is a legitimate answer. Writing it down is what makes it legitimate. The failure mode is neither: the row that reaches its ninth consecutive appearance while everyone has quietly stopped reading it.

Who should not buy a survey

Fifteen minutes a week does not sound like much until you know who spends it. The person running this meeting is the one the front desk escalates to, covering a missing member of staff and rebuilding tomorrow’s schedule at 4:40pm. A recurring commitment from a practice manager is not free, and this rhythm can decay for the same reason staff-effort review programmes do — capacity is the binding constraint, which is the argument I have made before about why the review ask cannot live at the front desk.

Our survey half is free, and you pay when a review posts, which means we would be paid for the review generation whether or not anyone ever opens the theme report. That is exactly why I want to say this plainly. If a practice is not going to give the rhythm fifteen minutes a week, it should not run a patient experience programme at all. Run the review generation on its own and be honest about what was bought. An unread survey is not neutral: it takes thirty seconds and a reply from a patient who has just been treated, and it carries an implied promise that somebody is going to do something with the answer.

The test of the programme is not whether the score went up. It is whether one named person, this week, changed one thing, and can tell you in a sentence what it was.

Sources

  1. Reichheld FF. “The One Number You Need to Grow.” Harvard Business Review, December 2003. hbr.org/2003/12/the-one-number-you-need-to-grow. Introduces the recommend question as the best single-question predictor of top-line growth, from two years of research linking survey answers to purchasing and referral behaviour.
  2. Bain & Company. “Introducing: the Net Promoter System.” Loyalty Insights. bain.com/insights/introducing-the-net-promoter-system-loyalty-insights. “Customers’ responses to the first question allow you to classify them as promoters (9–10), passives (7–8) or detractors (0–6).” The score is the percentage of promoters minus the percentage of detractors.
  3. Adams C, Walpola R, Schembri AM, Harrison R. “The ultimate question? Evaluating the use of Net Promoter Score in healthcare: a systematic review.” Health Expectations. 2022;25(5):2328–2339. doi:10.1111/hex.13577. onlinelibrary.wiley.com/doi/full/10.1111/hex.13577 (open access at pmc.ncbi.nlm.nih.gov/articles/PMC9615049). 12 studies met inclusion criteria; four determined NPS adds minimal value to healthcare improvement; free-text comment sections were favourably received in four; concludes NPS should be used in conjunction with a larger survey rather than standalone.
  4. Krol MW, de Boer D, Delnoij DM, Rademakers JJDJM. “The Net Promoter Score – an asset to patient experience surveys?” Health Expectations. 2015;18(6):3099–3109. doi:10.1111/hex.12297. onlinelibrary.wiley.com/doi/abs/10.1111/hex.12297. “The patient experiences from the surveys showed weaker associations with the NPS than with the global rating and the overall score.”
  5. Keiningham TL, Cooil B, Andreassen TW, Aksoy L. “A Longitudinal Examination of Net Promoter and Firm Revenue Growth.” Journal of Marketing. 2007;71(3):39–51. doi:10.1509/jmkg.71.3.039. journals.sagepub.com/doi/10.1509/jmkg.71.3.039. 21 firms and more than 15,500 interviews from the Norwegian Customer Satisfaction Barometer; fails to replicate the claimed superiority of Net Promoter over other measures in the industries cited as exemplars.

Want this kind of thinking applied to your practice?

Twenty minutes with us. We'll audit your current review velocity and tell you honestly whether applaud fits.