Skip to content
YibudYibud

Before a recommend score counts as evidence

NPS for Early-Stage Startups: What the Score Can and Can't Tell You

A Net Promoter Score is easy to collect and easy to misuse. Founders survey a waitlist, or twenty people who clicked a beta invite, then treat +42 as validation. The number only measures stated recommendation intent among people who already used you. It does not predict your startup. If you run it at all, write the tripwire first — and read behavior before you read the score.

Last updated: October 8, 2026

Direct answer

What can NPS tell an early-stage startup?

NPS asks how likely someone is to recommend you on a 0–10 scale. The score is the percent of promoters (9–10) minus the percent of detractors (0–6), so it runs from −100 to +100. It measures stated recommendation intent among people who already used you. For an early startup it is weak evidence: people who have not used a product cannot really answer it, and with a few dozen responses the score swings by tens of points. Use behavior — usage, payment, retention — first, and treat NPS, if you run it, as one pre-written tripwire.

Key takeaways

What to remember

  • NPS = % promoters (9–10) − % detractors (0–6). The range is −100 to +100. Passives (7–8) sit in the denominator and move neither side.
  • It is not the Sean Ellis product-market-fit question. Ellis asks how disappointed current users would be if they could no longer use the product. Do not invent a mapping between the two.
  • Asking a waitlist or landing-page visitor is asking about a hypothetical. That is opinion, not evidence.
  • At n = 30, a +10 NPS with a 40% promoters / 30% detractors mix can sit in a band about 60 points wide. Never report the score without n.
  • If you run NPS, survey recent real usage, always ask why, compare to your own earlier score, and write signal + threshold + window + action before the send.

The instrument

What NPS actually measures

Bain & Company’s public explainer, “Measuring Your Net Promoter Score,” writes the question as: How likely are you to recommend us to a friend or colleague? Answers sit on a 0–10 scale. The three groups below are Bain’s scoring rule, not a Yibud invention. The score is a stated-intent metric among people who already have a relationship with you. It is not usage, not payment, and not retention.

NPS respondent groups on the 0–10 scale
GroupScoresHow they enter the score
Promoters9–10Counted in % promoters. They raise NPS.
Passives7–8Neither side. They sit in the sample and dilute both percentages.
Detractors0–6Counted in % detractors. They lower NPS.

The arithmetic

NPS = (% answering 9–10) − (% answering 0–6)

A sample that is 40% promoters and 30% detractors scores +10. A sample that is all 8s scores 0. The range is −100 to +100. Bain’s system also pairs the score with an open “why” follow-up. Dropping the why is how a number becomes theater.

Not the same question

NPS is not the Sean Ellis product-market-fit survey

Founders collapse these into one “loyalty score.” They are different questions with different threshold logic. The Ellis wording below is reused exactly from Yibud’s product-market-fit survey page. This page invents no mapping, no conversion table, and no “NPS of X equals 40% very disappointed.”

NPS versus the Sean Ellis product-market-fit question
FieldNPSSean Ellis PMF survey
QuestionHow likely are you to recommend us to a friend or colleague? (0–10)How would you feel if you could no longer use [ProductName]?
Who you askPeople who already used you. A waitlist cannot really answer.People with recent real usage — Ellis’s example: took a ride, not just signed up. Exclude N/A.
Score% promoters (9–10) minus % detractors (0–6). Range −100 to +100.Percent very disappointed among valid current users. Exclude N/A. Somewhat disappointed does not fire the tripwire.
Threshold logicThis page publishes no “good NPS” bar and no industry average.Ellis’s experience: around 40 percent very disappointed is when sustainable growth becomes possible. Treat 40 percent as a default you may adopt or rewrite before the send — not a law.

A high NPS is not product-market fit. A high very-disappointed percent is not an NPS. Do not average them. Do not swap the questions and keep the same bar.

Sean Ellis PMF survey — write the tripwire first →

Why it is weak early

Three reasons an early NPS is thin evidence

A common pattern: twenty warm answers get treated as a market. The score looks decisive because it is a single number. But the sample is often people who have never done the job, plus friends who would say 9 to be kind. The number does not fail the founder. The question and the sample do.

1. Hypothetical respondents cannot recommend a product they have not used

A waitlist, a landing-page visitor, or someone who only saw a mockup is answering a story. That is opinion. Money-moving evidence lives on the willingness-to-pay page. Reach and intent before usage live on the smoke-test page. NPS starts after someone has actually used you.

2. Who you survey is the result

Friends, advisors, and people who owe you a reply inflate promoters. People who churned and will not open the email are missing detractors. Bain’s scoring rule assumes a current-customer sample. Pad the list to chase a nicer number and you measured courtesy.

3. A few dozen answers swing the score by tens of points

The next section is Yibud’s own arithmetic, labeled as a normal-approximation estimate — not a law. At n = 30, a +10 can sit in a band that runs from about −20 to +40. That is not a precision instrument. It is weather you should not bet a sprint on.

Small-sample noise

Why +10 at n = 30 is not a number you can steer by

Treat each answer as +1 (promoter), 0 (passive), or −1 (detractor). The Net Promoter Score, written as a proportion from −1 to +1, then has this standard error — a normal-approximation estimate, not a statistical law:

Standard error of NPS (normal approximation)

SE = sqrt( (p_promoters + p_detractors − (p_promoters − p_detractors)²) / n )

p_promoters and p_detractors are sample proportions. A 95% margin is about 1.96 × SE, then multiplied by 100 to read as NPS points. This is an approximation. It is not a confidence-interval theorem for your category, and it is not a Yibud engine rule.

Worked example at 40% promoters and 30% detractors, NPS = +10
nSE (proportion)95% marginAround +10
30≈ 0.152≈ ±30 points (≈ ±29.7)about −20 to +40
100≈ 0.083≈ ±16 points (≈ ±16.3)about −6 to +26

Worked mix, used exactly: 40% promoters, 30% detractors → NPS = +10. Variance term: 0.70 − 0.01 = 0.69. At n = 30, SE ≈ 0.152. At n = 100, SE ≈ 0.083. The true score at n = 30 could plausibly sit anywhere from about −20 to +40. That is why this page refuses industry “good NPS” charts: the chart is tighter than your sample. The general sample-size idea — read a rate as a range, write the rule before you look — lives on the smoke-test sample-size page.

How many visitors a smoke test needs →

The research debate

Reichheld proposed a growth metric. Keiningham et al. did not replicate the claimed superiority.

State this carefully. One article introduced the recommend question as a growth metric. One later paper, using a different dataset, failed to replicate the claimed “clear superiority” of Net Promoter over other satisfaction measures. That is not the same as proving NPS is useless. This page cites only those two pieces of research.

Reichheld (2003) — the article that introduced the recommend question as a growth metric

Frederick F. Reichheld, “The One Number You Need to Grow,” Harvard Business Review, December 2003 (reprint R0312C). Cite this only as the article that put the recommend question forward as a growth metric. This page does not quote it and does not cite numbers from it — the article is paywalled and those figures were not verified for this write-up.

Keiningham, Cooil, Andreassen, Aksoy (2007) — failed to replicate the claimed superiority

Timothy L. Keiningham, Bruce Cooil, Tor Wallin Andreassen, and Lerzan Aksoy, “A Longitudinal Examination of Net Promoter and Firm Revenue Growth,” Journal of Marketing 71(3), July 2007, pp. 39–51 (DOI 10.1509/jmkg.71.3.039). They used longitudinal data from 21 firms and 15,500-plus interviews from the Norwegian Customer Satisfaction Barometer, compared Net Promoter with measures such as the American Customer Satisfaction Index (ACSI), and could not replicate the claimed “clear superiority” of Net Promoter in the industries Reichheld cites as exemplars. The paper received the 2007 MSI / H. Paul Root Award. Phrase it as failed to replicate the claimed superiority — not “proved NPS is useless.”

For an early-stage founder the debate is almost beside the point. You do not have 21 firms and 15,500 interviews. You have a few dozen people, many of whom have not used the product. The instrument’s owner and its critics were arguing about established firms. Your problem is sample, timing, and whether anyone has done the job.

If you still run it

How to use NPS as a tripwire — not as a success score

This is Yibud’s suggested template, not a law. Run it only after recent real usage exists. Write the four fields before the first send. Honor them on the end date. If you cannot name real usage, this is the wrong instrument.

  1. 1

    Survey only people with real recent usage

    Name the event that counts — published a digest a teammate opened, closed a job the customer accepted, ran the weekly report for a live account. Last two weeks is a usable window. Signups, waitlist members, and “saw the landing page” are not a sample.

  2. 2

    Always ask the open “why” follow-up

    Bain’s system pairs the score with a reason. The number without the reason is a mood. The follow-up is how you learn what promoters would actually tell a colleague, and what detractors hit this week.

  3. 3

    Write signal + threshold + window + action before you survey

    Example you may adapt, not a law: “If, among ≥40 users who used it in the last 2 weeks, detractors outnumber promoters, we stop adding features and interview 5 detractors this week.” Date the page. A kinder bar after the results is a recap.

  4. 4

    Compare to your own earlier score, not to a published industry benchmark

    This page publishes no “SaaS average” and no “good NPS is 50+.” Those charts mix companies, methods, and sample sizes you cannot see. Your last score, from the same question, the same usage filter, and a reported n, is the only comparison that belongs on the page.

  5. 5

    Never report NPS without n

    Write “+10 (n = 30)” or do not write the number. A score without a sample size invites someone — you, a co-founder, a future you — to treat weather as a trend.

What you leave with

The NPS tripwire (copy this)

If the sitting produced a mood, you did not finish. Fill the brackets. Leave nothing as “users,” “soon,” or “we’ll know it when we see it.”

The fill-in contract

Product: [name]. Real usage (must have happened): [one event — not signup]. Recent window: last [2 weeks / your dates]. Question: How likely are you to recommend us to a friend or colleague? (0–10) + open why. Signal: NPS = % promoters (9–10) − % detractors (0–6), reported with n. Threshold: [e.g. detractors outnumber promoters / NPS below our last score of X at n = Y]. Sample you will honor: [e.g. ≥40 users with recent usage]. Window: [start]–[end]. Action if miss: [stop adding features / interview 5 detractors this week / run a named cheaper test].

The “≥40 users / detractors outnumber promoters / interview 5 detractors” line above is Yibud’s suggested template, not a published law and not a Bain rule. Rewrite the numbers before the send if you have a reason. Do not rewrite them after.

Example tripwire (Yibud's suggested template, not a law)

If, among ≥40 users who used it in the last 2 weeks, detractors outnumber promoters → stop adding features the same day and interview 5 detractors this week. Report the score as NPS (n). Compare only to our last score from the same filter. Do not add waitlist names to chase n.

Passives still count in n. Friends who never did the named event do not. A Yibud report score does not fire this tripwire.

Not the same page

NPS vs the Ellis survey, sample-size math, behavior, interviews, Yibud scores, and kill lines

Several Yibud pages sit next to this one. Mixing them up turns a recommend score, a waitlist, or a mid-band engine number into “the market loves us.”

Sean Ellis PMF survey — disappointment among current users, not recommend intent

That page’s signal is the percent of recent real users who would be very disappointed if the product disappeared. This page’s signal is % promoters minus % detractors. Different question. Different threshold logic. No mapping.

Product-market fit survey →

Smoke-test sample size — general sample-size logic, not NPS noise

That page reads k signups out of n qualified visitors as a 95% Wilson interval. This page applies a normal-approximation standard error to NPS. A conversion interval is not an NPS band. An NPS band is not a signup rate.

How many visitors a smoke test needs →

Willingness to pay and smoke tests — behavior beats stated intent

A deposit, pre-order, paid concierge, or checkout is money moving. A smoke-test click is reach and intent before usage. NPS is what people say they would tell a friend. Ask the money or the click first when that is the unknown. Do not treat a 9 as a charge.

Interview count and the Mom Test — “why” beats a score at tiny n

At a few dozen answers, five detractor interviews usually teach more than the point estimate. Saturation is coverage of needs in one segment. The Mom Test is last-week behavior with no pitch. Neither is an NPS.

Yibud scores — not NPS, and not a success prediction

The validation-score and score-calculator pages explain Startup MRI numbers from a deterministic rule engine on a one-line idea plus five questions. Those scores are flashlights on the weakest assumption. They are not Net Promoter. Do not read a mid-band Yibud score as +40 NPS, and do not read +40 as a Yibud GO.

Kill criteria and critical assumptions — where a tripwire becomes a decision

Those pages write Stop, or name the one claim that would kill the idea if false. This page can supply one optional input those contracts later read: a dated NPS against a bar you wrote first, among people with recent usage, reported with n. It is not the decision sitting.

Worked example

One NPS week you can copy (illustrative — fictional)

The week below is illustrative. Names, product, and counts are invented so you can see the setup. It is not a study, not an industry benchmark, and not a Yibud report. Copy the method, not the numbers.

The recommend bet, written before the first send

Illustrative names: founder Mara Chen at Shiftlog, a tool that turns a warehouse shift handover into a shared log. Real usage, written before the send: “published a handover that the next shift opened.” Signup and “connected Slack” do not count. The tempting next step is a feature sprint because a waitlist of 80 “would recommend it.” The current bet assumes supervisors who already published a handover would recommend Shiftlog to a peer.

The tripwire, written the same sitting

Signal: NPS among users who published a handover in the last 2 weeks, reported with n, plus the open why. Threshold: if detractors outnumber promoters, or if NPS falls below the last score from the same filter. Sample: honor the read only at ≥40 such users. Window: first send through day 14. Action if miss: stop the feature sprint and interview 5 detractors this week. Waitlist names do not enter n. If n is still under 40 on day 14, that also counts as a miss: the feature sprint does not start on a recommend story, and the action still runs.

What a miss looks like on day 14

Illustrative close: 28 people with recent real usage. 11 promoters, 8 passives, 9 detractors. NPS = (11/28 − 9/28) × 100 ≈ +7, n = 28. The waitlist-only batch, which should never have been scored, read +44. Day 14 is a miss under the sample rule you wrote: n never reached 40. Detractors did not outnumber promoters (9 vs 11), but at n = 28 that two-person gap is noise, not a pass. The feature sprint stops, and the why notes from the nine detractors are the week's work. +7 is not “almost a good score.” It is a small-n number you already said you would not steer by.

What must not happen in the window

Do not add waitlist emails to chase 40. Do not drop the why. Do not compare +7 to a “SaaS average” this page does not publish. Do not treat +7 as product-market fit, and do not treat a Yibud score as this NPS. Do not ask “would you pay?” and file the yes under this tripwire — that is the willingness-to-pay page.

Write the tripwire before the first send. If day 14 arrives and the sample or the mix misses the bar you wrote, honor the action. If it cleared the bar, Continue to the next cheapest test — usually money or a named usage metric, not a factory. The method is the dated page, not the story you tell after.

Where Yibud fits

Neither NPS nor a Yibud score is a success prediction

Yibud is a free, no-signup startup idea validator. Scores come from a deterministic rule engine, not from a language model guessing success. Optional AI text, when it is used, only polishes prose. It does not invent the number. A Startup MRI report can name a weak “people will tell a friend” or distribution assumption, then a Day-7 Continue / Refine / Re-test / Stop signal. NPS work, if you do it at all, sits next to that assumption — write the tripwire, do not hope a mid-band score means “we already have promoters.” If you already have a report, start with the named assumption, not the overall score. Then run the window yourself. The method on this page stands alone if you never open the analyzer.

Analyze my idea →

Common mistakes

What usually wastes an early NPS

Surveying a waitlist or landing-page visitors

They have not used you. A 9 about a landing page is a compliment to copy. Ask willingness to pay or run a smoke test if the unknown is demand. Come back here after a named usage event.

Reporting NPS without n

“We’re at +32” is not a sentence. “+32 (n = 19)” is. At n = 30 the same mix that produces +10 can sit near −20 or +40. Hide n and you hide the weather.

Comparing to a random industry benchmark

This page publishes no “SaaS average is 30” and no “good NPS is 50+.” Those charts are not a sample you can join. Compare to your last score from the same question and the same usage filter.

Dropping the “why” follow-up

Bain’s system pairs the score with a reason. A promoter who cannot name what they would tell a colleague, and a detractor you never interview, leave you with a number you cannot act on.

Treating a high NPS as product-market fit — or as a forecast

Recommend intent is not “very disappointed if it disappeared.” It is not payment. It is not a Yibud GO. Neither NPS nor a Yibud report score predicts that the startup will succeed. A tripwire you dated before the send is a decision aid. A trophy is not.

Sources

Where these ideas come from

Net Promoter, Net Promoter Score, and NPS are trademarks of Bain & Company, Inc., Fred Reichheld, and Satmetrix Systems, Inc.

In one paragraph

Summary you can quote

NPS asks how likely someone is to recommend you on a 0–10 scale. The score is % promoters (9–10) minus % detractors (0–6) and runs from −100 to +100. It measures stated recommendation intent among people who already used you. It is not the Sean Ellis product-market-fit question, which asks how disappointed current users would be if they could no longer use the product — around 40 percent very disappointed is Ellis’s experience bar, not a mapping to NPS. Asking a waitlist is asking a hypothetical. A normal-approximation estimate on a 40% / 30% mix (NPS = +10) gives SE ≈ 0.152 at n = 30 (about ±30 points, so about −20 to +40) and SE ≈ 0.083 at n = 100 (about ±16 points). Reichheld (HBR, December 2003) introduced the recommend question as a growth metric; Keiningham et al. (Journal of Marketing, 2007; 21 firms, 15,500-plus interviews) failed to replicate the claimed “clear superiority” over measures such as the ACSI. If you run NPS, survey recent real usage, ask why, write signal + threshold + window + action first, compare only to your last score, and never report the number without n. Neither NPS nor a Yibud score predicts success.

FAQ

Questions founders actually ask

What is NPS?

Net Promoter Score asks: How likely are you to recommend us to a friend or colleague? on a 0–10 scale. Promoters are 9–10, passives are 7–8, detractors are 0–6. NPS = % promoters − % detractors, from −100 to +100. It measures stated recommend intent among people who already used you.

Is NPS the same as the Sean Ellis product-market-fit survey?

No. Ellis asks how you would feel if you could no longer use the product, and scores the percent very disappointed among valid current users (exclude N/A). NPS asks about recommending. Different question, different threshold logic. This page invents no mapping between them.

Can I ask NPS of a waitlist or of people who only saw a landing page?

You can ask. You cannot treat the answer as evidence. They have not used you, so they are scoring a hypothetical. That is opinion. Use a smoke test or a willingness-to-pay test for demand; come back to NPS after recent real usage.

How many NPS responses do I need?

Enough that you will not steer by weather. Yibud’s normal-approximation estimate, not a law: at 40% promoters / 30% detractors (NPS = +10), n = 30 has a 95% margin of about ±30 points; n = 100 about ±16. The suggested tripwire on this page honors a read among ≥40 recent users — rewrite that number before the send if you have a reason. Always report n.

Does a high NPS mean product-market fit? Is a Yibud score the same thing?

No and no. A high NPS is stated recommend intent, not “very disappointed if it disappeared,” not payment, and not a forecast. Yibud scores come from a deterministic rule engine on your idea inputs. They point at the weakest assumption to test first. Neither number predicts success. You do not need an account to use this page.

Write the recommend tripwire before the next sprint

Yibud scores weak assumptions with a deterministic engine. You write signal, threshold, window, and action while you are still honest — then an NPS, if you collect one, is a tripwire, not a trophy.

Generate a Yibud report →