Before a count counts as coverage
How Many Customer Interviews Is Enough — Write the Stopping Rule Before You Count
A headcount is not a verdict. “How many customer interviews is enough?” is a coverage question in one tightly defined segment — have we heard most of the needs and pains — not a validation question about whether anyone will pay or use the product. Write the stopping rule before the first call. Then the number is a contract you honor, not a trophy you invent on Sunday.
Last updated: October 5, 2026
Direct answer
How many customer interviews is enough?
Interview count answers a coverage question in one segment, not a demand question. Write a stopping rule before the first call — segment, batch size, minimum, saturation check, cap, and the named action — then interview in batches of about 5–6 and stop when a full batch adds no new important need or pain code. Hennink and Kaiser’s 2022 review of empirical tests found saturation at 9–17 interviews in homogeneous, narrowly aimed studies; Guest, Bunce and Johnson (2006) saw saturation within the first 12 of 60 interviews; Griffin and Hauser (1993) hypothesized 20–30 one-on-ones for 90–95 percent of needs in their product category, per segment — if you pass that band and still hear new needs, split the segment rather than keep going.
Key takeaways
What to remember
- Count measures coverage of needs in one segment. It does not measure demand, payment, or product-market fit.
- Write the stopping rule before interview 1: segment, batch size, minimum N, saturation check, cap, and the named action when the rule fires.
- A starting rule drawn from the studies below — not a law: batches of about 5–6; stop when a full batch adds no new important code; expect roughly 9–17 in a homogeneous segment.
- Heterogeneous ideas need a separate count per segment. A new segment beats a 40th interview in the old one.
- Saturation means no new themes. It does not mean the idea is validated. After saturation, move to behavior or money evidence.
Why it matters
A magic number is how founders keep booking calls instead of deciding
I have watched founders treat “we need one more interview” as a personality trait. Two compliments become ten. Ten mixed jobs become a round number they saw in a thread. The count was never going to tell them whether anyone would pay. It could only tell them whether they had heard most of the needs in one group of people who share a job. Guest, Bunce and Johnson (2006) ran 60 in-depth interviews with women in Ghana and Nigeria and found saturation within the first 12; they caution that 12 fits relatively homogeneous groups and focused aims. Hennink and Kaiser (2022) reviewed 23 articles and found empirical studies reaching saturation at 9–17 interviews — again, mainly homogeneous populations with narrow objectives. Griffin and Hauser (1993) interviewed 30 customers in one product category and hypothesized 20–30 one-on-ones for 90–95 percent of customer needs, per segment. Even the 30 customers in that study did not capture every need; their model estimated those 30 customers accounted for 89.8 percent of all needs, because low-probability needs were still missing. A rule written after the compliments is a recap.
Vocabulary
What the six stopping-rule fields mean
The page is not “keep talking until it feels done.” It is a dated contract: who counts as one segment, how you batch, what “no new codes” means, where you cap, and what you do when the rule fires. Each field needs a log a stranger could read.
1. Segment — one job, one context, one buyer
A segment is a group that shares the same recent job and constraint — not “SMBs,” not “people who might like this.” Guest, Bunce and Johnson caution that their 12-interview finding fits relatively homogeneous groups. Mix two jobs and you are counting coverage you do not have.
2. Saturation — a full batch adds no new important code
Saturation here means no new important need or pain codes in a complete batch. It does not mean people liked the idea. It does not mean they will pay. Hennink and Kaiser are explicit: their 9–17 range is not a generic rule, and outliers (multi-country work, meta-themes, code-meaning saturation) needed larger samples.
3. Batch — about 5–6 talks, then you recode
A batch is a small set you finish before you decide whether to continue. Guest, Bunce and Johnson saw basic elements of metathemes as early as 6 interviews. Treat about 5–6 as a starting batch size derived from that early-theme finding and from the practical need to recode, not as a law.
4. Stopping rule — six fields, written before call 1
Segment, batch size, minimum N, saturation check, cap, and the named action. If any field is blank, you do not have a rule. A number invented after a warm week is a recap.
5. Coverage vs validation
Coverage asks: have we heard most of the needs and pains in this segment? Validation asks: will they pay, use, or commit? Interview count measures the first. Deposits, checkouts, smoke-test clicks, and scarce commitments measure the second. Do not convert one into the other.
The evidence
What the four studies actually measured — and what they did not
These are study findings and authors’ hypotheses, labeled as such. They are not a personal forecast and not a Yibud rule. Do not import numbers this page does not publish.
Guest, Bunce & Johnson (2006) — saturation within 12 of 60
Sixty in-depth interviews with women in Ghana and Nigeria. Saturation occurred within the first 12 interviews. Basic elements of metathemes were present as early as 6. After 12 interviews they had 92 percent (100 of 109) of the codes for the 30 Ghana transcripts and 88 percent (100 of 114) of codes across all 60. The authors caution that 12 fits relatively homogeneous groups and focused aims.
Hennink & Kaiser (2022) — empirical saturation at 9–17
A systematic review of 23 articles (17 empirical, 6 statistical modeling). Empirical studies reached saturation at 9–17 interviews or 4–8 focus groups, mainly with homogeneous populations and narrowly defined objectives. Outliers needed larger samples. The authors do not offer a generic rule.
Griffin & Hauser (1993) — 20 interviews, over 90 percent of needs found by 30
One-on-one interviews with 30 customers in one product category. About 20 interviews surfaced over 90 percent of the needs found by 30. The authors hypothesize that 20–30 interviews are needed for 90–95 percent of customer needs, and that one-on-ones may be more cost-effective than focus groups. They also say multiple analysts should read transcripts. Label these as their study and hypotheses, per segment — not a universal law.
Nielsen (2000) — contrast only: 5 users for usability, not discovery
Jakob Nielsen’s “Why You Only Need to Test with 5 Users” is about usability testing of an existing design. About 85 percent of usability problems with 5 users; better to run several small rounds. It is not a discovery-interview sample size. Borrowing “5” for customer interviews is a common misuse.
A starting rule you may adopt — written before call 1
Interview in batches of about 5–6 within one tightly defined segment. After each batch, list new need and pain codes. Stop when a full batch adds no new important code. Expect this to land roughly in the 9–17 range for a homogeneous segment (Hennink & Kaiser). If you pass about 20–30 in one segment and still hear new needs (Griffin & Hauser’s hypothesized band), the segment is probably too broad — split it rather than keep going.
Not the same page
Interview count vs scripts, PMF surveys, money tests, who-filters, contracts, and scores
Several Yibud pages sit next to this one. Mixing them up turns a polite script, a usage percent, or a mid-band engine score into “we did enough interviews.”
Mom Test, discovery, solution interview — how to ask, not how many
Those pages are conversations: last-week behavior with no pitch, the discovery process, or a scripted artifact show after the problem is evidenced. This page is how many talks in one segment and when to stop. A good question is not a saturation check. A completed walkthrough is not coverage.
Product-market fit survey — a usage percent, not saturation
That page’s signal is the percent of recent real users who would be very disappointed if the product disappeared. This page’s signal is whether a batch in one segment still adds new need codes. Saturation is not a Sean Ellis percent. A very-disappointed score is not interview coverage.
Product-market fit survey →WTP, smoke, fake-door, landing page — behavior and money after coverage
Those pages measure reach, clicks, deposits, or checkouts. They usually sit after interviews saturate — or instead of more interviews when the expensive unknown is already money or reach. A waitlist is not a new need code. A deposit is not saturation.
Earlyvangelist and design partner — who, and scarce commitment
Those pages filter a named person (Blank’s five characteristics) or a scarce co-dev commitment. This page counts coverage inside a segment you already named. Matching the five is not saturation. A signed partner is not a batch of codes.
Kill, pivot, assumption, decision — contracts that can later read this rule
Those pages write Stop, rewrite a dimension, name one claim, or read a week of evidence you already have. This page supplies one input those contracts can later read: a dated coverage check in one segment. It is not a Stop / Pivot / Continue sitting.
Yibud scores — not a saturation measure
The validation-score and score-calculator pages explain Startup MRI numbers from a deterministic engine on a one-line idea plus five questions. Those scores are flashlights on assumptions. They are not interview saturation and not a Guest-or-Hennink count. Do not read a mid-band score as “enough interviews,” and do not read 12 talks as a Yibud GO.
The sitting
How to write and run the interview stopping rule in one sitting
Do this before you book the next stranger. Set a timer. Leave the calendar closed until the six fields are filled. If you have a co-founder, each of you writes a draft alone, then you keep the tighter segment.
- 1
Write the six fields before anyone is invited
Segment: one job, one context. Batch size: about 5–6 as a starting size. Minimum N: at least one full batch. Saturation check: a full batch adds no new important need or pain code. Cap: the point where you split instead of continue — Griffin and Hauser’s hypothesized 20–30 band is a default you may adopt, labeled as their hypothesis. Action: stop this segment, split it, or move to a named money or behavior test. If any field is blank, you are not ready to book.
- 2
Define the segment so a stranger could recruit it
Write one sentence: the recent job, the constraint, the buyer. “Independent US bookkeepers who chase AR by email for 5–20 clients” is a segment. “Small businesses” is a conference. Griffin and Hauser: even the 30 customers in that study did not capture every need; their model estimated those 30 customers accounted for 89.8 percent of all needs, because low-probability needs were still missing. A new segment beats a 40th interview in the old one.
- 3
Interview in batches of about 5–6, then recode
Finish the batch. List every new need and pain code before you book the next person. Guest, Bunce and Johnson saw basic metatheme elements as early as 6. Do not average two jobs into one “insight” because the week is busy.
- 4
Apply the saturation check you wrote
If the completed batch added no new important code, the rule fired. Stop this segment. Hennink and Kaiser’s 9–17 range is where many homogeneous, narrow empirical studies landed — treat it as a place you may land, not a target to hit.
- 5
Honor the cap or split — do not invent a kinder band
If you pass about 20–30 in one segment and still hear new important codes, split the segment. That is Griffin and Hauser’s hypothesized 90–95 percent band, labeled as theirs, for one product category, per segment. Stretching the cap because the stories were warm is a new page with a reason.
- 6
Do the named action the same day the rule fires
Saturation means no new themes, not “the idea is validated.” The action is usually: stop interviewing this segment and run a named willingness-to-pay, smoke, fake-door, landing-page, design-partner, or earlyvangelist test. Compliments do not fire a different action.
What you leave with
The interview stopping-rule contract (copy this)
If the sitting produced a mood, you did not finish. The page should have one filled rule a stranger could apply. Fill the brackets. Leave nothing as “users,” “soon,” or “we’ll know it when we see it.”
The stopping rule
Segment: [one job + one context + one buyer — not “SMBs”]. Batch size: [about 5–6 starting size]. Minimum N: [at least one full batch]. Saturation check: a full batch adds no new important need/pain code. Cap: [~20–30 in one segment → split; Griffin & Hauser hypothesis, one category]. Action when the rule fires: [stop this segment / split it / run a named money or behavior test]. Codes you will log: [need and pain phrases in their words].
If any bracket still says “customers,” “enough,” or “later,” the contract is not written. Interviews from two jobs do not share a count. Compliments are not codes.
What does not count as the signal
Not this signal: a Yibud score, a compliment, “I’d use that,” a Sean Ellis very-disappointed percent, a waitlist, a smoke or fake-door click, a deposit or checkout, matching Blank’s five, a signed design partner, a completed solution-interview walkthrough, Nielsen’s 5 usability users. Optional later: those instruments have their own pages.
You may still take notes on praise or money. You may not let them fire this stopping rule. New important need/pain codes in a completed batch fire this rule — or the absence of them does.
Example starting rule (template — the counts are the research defaults, labeled)
In [named segment], interview in batches of 6. After each batch, list new need/pain codes. Minimum: 6. Stop when a full batch adds no new important code (expect roughly 9–17 in a homogeneous segment — Hennink & Kaiser 2022, not a law). If you pass ~20–30 and still hear new needs, split the segment (Griffin & Hauser 1993 hypothesis, one category, per segment). Action: stop this segment and run a named WTP or smoke test.
9–17, 12, and 20–30 are from the studies named on this page, labeled as findings or hypotheses. They are not Yibud results and not a claim about what usually happens in your category. This page invents no other N.
The fork
When interview coverage is the expensive unknown — and when it is not
This page decides only whether “have we heard most of the needs in this one segment?” is the unknown you should buy evidence for now. How you ask, money, scarce commitment, reach, usage surveys, and decision contracts have their own pages. Do not keep interviewing a segment you already saturated.
Write the stopping rule first
You can name one segment. The expensive unknown is coverage of needs and pains. Write the six fields. Book the first batch. Do not start with ads or a factory.
Move to money or behavior first
A full batch already added no new important code, or the unknown is already whether they will pay or click. Use the WTP, smoke, fake-door, or landing-page page. More interviews will not move money.
Split the segment first
You passed about 20–30 in what you called one segment and still hear new important codes. The segment is probably too broad. A new segment beats a 40th interview in the old one.
When an interview stopping rule is the expensive unknown
You can name one job and one buyer. You do not yet know whether the next batch will still add new need codes. You have not written the cap or the action. That is this page. A weak “we don’t know the customer” line on a Yibud report sits near this fork — as a flashlight on which segment or assumption to interview about, not as a count.
When an interview stopping rule is the wrong next test
You cannot name a segment. You are still asking “would you use this?” You only needed a script, a usage survey of people who already did the real thing, a scarce-commitment partner, or a charge. You are mixing three jobs into one count. You are running a usability test and quoting Nielsen’s 5. Run those pages instead. Counting compliments across mixed segments is theater.
Worked example
One stopping-rule contract you can copy (illustrative)
The week below is illustrative — names, product, and counts are invented so you can see the setup. It is not a study, not a saturation benchmark, and not a Yibud report. Copy the method, not the numbers. Guest’s 12, Hennink and Kaiser’s 9–17, and Griffin and Hauser’s 20–30 stay in their papers; they are not this story.
The coverage bet, written before the first call
Illustrative names: founder Mei Chen at Ledgerping, a B2B SaaS that drafts payment-chase emails for independent bookkeepers. Segment, written before call 1: US independent bookkeepers who currently chase accounts receivable by email for 5–20 clients. “Accountants” and “SMBs” do not count. The tempting next step is a 40-interview sprint across anyone who will take a calendar hold.
The stopping rule, written the same sitting
Segment: as above. Batch size: 6. Minimum: one full batch. Saturation check: a completed batch adds no new important need or pain code. Cap: 24 in this segment — if new important codes continue, split (Griffin and Hauser’s 20–30 hypothesis, adopted as a cap, labeled as theirs). Action if the rule fires: stop interviewing this segment and run a named deposit test on the two recurring pains. Compliments do not count as codes.
What a stop looks like after three batches
Illustrative close: batch 1 (talks 1–6) adds many codes — Friday-night chasing, clients who will not open a portal, shame about late invoices. Batch 2 (7–12) adds two codes. Batch 3 (13–18) adds none that the founder marked important. That is a coverage stop you already named — move to money evidence. It is not a reason to book twenty more “because people were warm.” Eighteen talks in this fiction is not Guest’s 12 and not a claim about your week.
What must not happen in the window
Do not fold agency bookkeepers and in-house controllers into the same count. Do not treat “I’d use that” as a new need code. Do not borrow Nielsen’s 5 and stop after five compliments. Do not treat a Yibud score as saturation. Do not keep going past the cap because the stories rhyme — that is a split, or a new page with a reason. If you cannot find one batch in the named segment, you may have a reach problem — discovery or smoke, not a fake pass.
Write the stopping rule before the first call. If a full batch adds no new important code, honor the action. If you pass the cap and still hear new needs, split the segment. After saturation, Continue to the next cheapest test — often money or a design-partner ask, not a factory. The method is the dated page, not the story you tell after.
Where Yibud fits
Use the free validator as a flashlight, then write the stopping rule yourself
Yibud is a free, no-signup startup idea validator. Scores come from a deterministic rule engine, not from a language model guessing success. Optional AI text, when it is used, only polishes prose. It does not invent the number. A Startup MRI report can name a weak customer or assumption dimension, then a Day-7 Continue / Refine / Re-test / Stop signal. Interview-count work sits next to that assumption — use the weakest dimension to pick which segment or claim to interview about. Do not hope a mid-band score means “we already have enough interviews.” If you already have a report, start the contract with the named assumption — not with the overall score. Then run the batches yourself. The method stands alone if you never open the analyzer.
Analyze my idea →Common mistakes
What usually wastes the interview stopping rule
Misusing Nielsen’s “5 users” as a discovery sample
Nielsen (2000) wrote about usability testing of an existing design: about 85 percent of usability problems with 5 users, and several small rounds beat one big one. That is not a discovery-interview count. Five compliments about a slide are not saturation.
Counting interviews across mixed segments
Twelve talks with three different jobs is not Guest’s 12. Coverage is per segment. Griffin and Hauser: even the 30 customers in that study did not capture every need; their model estimated those 30 customers accounted for 89.8 percent of all needs, because low-probability needs were still missing. A new segment beats a 40th interview in the old one.
Treating saturation as validation
No new themes means you have probably heard the needs in that group. It does not mean they will pay, click, or stay. After the rule fires, move to a named money or behavior test. Do not announce product-market fit.
Stopping on compliments — or never writing the action
“I’d use that” is politeness. It is not a new code and not a reason to stop early. If you did not name the action before call 1, every warm talk becomes “qualitative insight.” Write the fail line on the same page as the segment.
Treating 12 or 20–30 as a law or a trophy
Guest, Bunce and Johnson caution that 12 fits homogeneous groups and focused aims. Hennink and Kaiser say 9–17 is not a generic rule. Griffin and Hauser’s 20–30 is a hypothesis from one product category, per segment. Adopt the starting rule or rewrite it before the first call. A band invented after a messy week is a recap.
Sources
Where these ideas come from
- Guest, Bunce & Johnson, “How Many Interviews Are Enough? An Experiment with Data Saturation and Variability,” Field Methods 18(1):59–82 (2006) — Used for 60 in-depth interviews with women in Ghana and Nigeria; saturation within the first 12 interviews; basic elements of metathemes present as early as 6; after 12 interviews, 92 percent (100 of 109) of codes for the 30 Ghana transcripts and 88 percent (100 of 114) of codes across all 60; and the authors’ caution that 12 fits relatively homogeneous groups and focused aims. This page does not invent other interview counts.
- Hennink & Kaiser, “Sample sizes for saturation in qualitative research: A systematic review of empirical tests,” Social Science & Medicine 292:114523 (2022) — Used for the review of 23 articles (17 empirical, 6 statistical modeling); empirical saturation at 9–17 interviews or 4–8 focus groups, mainly with homogeneous populations and narrowly defined objectives; outliers (multi-country, meta-themes, code-meaning saturation) needing larger samples; and the authors’ statement that this is not a generic rule. PubMed record: https://pubmed.ncbi.nlm.nih.gov/34785096/
- Griffin & Hauser, “The Voice of the Customer,” Marketing Science 12(1):1–27 (1993) — Used for 30 one-on-one interviews; about 20 interviews surfacing over 90 percent of the needs found by 30; the authors’ hypothesis that 20–30 interviews are needed for 90–95 percent of customer needs; the hypothesis that one-on-ones may be more cost-effective than focus groups; and the note that multiple analysts should read transcripts. Labeled as their study and hypotheses (one product category, per segment), not a universal law. MIT Sloan PDF
- Jakob Nielsen, “Why You Only Need to Test with 5 Users,” Nielsen Norman Group (2000) — Used only as contrast: usability testing of a design; about 85 percent of usability problems with 5 users; better to run several small rounds. Not a discovery-interview sample size.
In one paragraph
Summary you can quote
How many customer interviews is enough is a coverage question in one segment, not a demand question. Write a stopping rule before the first call: segment, batch size, minimum N, saturation check, cap, and the named action. A starting rule derived from the research on this page — not a law — is to interview in batches of about 5–6 and stop when a full batch adds no new important need or pain code; expect roughly 9–17 in a homogeneous, narrow study (Hennink & Kaiser 2022); Guest, Bunce and Johnson (2006) saw saturation within the first 12 of 60 interviews and caution that 12 fits homogeneous groups and focused aims; Griffin and Hauser (1993) hypothesized 20–30 one-on-ones for 90–95 percent of needs in their category, per segment — if you pass that band and still hear new needs, split the segment. Nielsen’s 5 users is usability testing of a design (~85 percent of usability problems; several small rounds), not discovery. Saturation means no new themes, not that the idea is validated. After the rule fires, move to behavior or money evidence. Yibud scores are not saturation measures. Matching the rule you wrote first is the method.
FAQ
Questions founders actually ask
How many customer interviews is enough?
Enough for coverage in one segment: write a stopping rule, interview in batches of about 5–6, and stop when a full batch adds no new important need or pain code. Hennink and Kaiser (2022) found empirical saturation at 9–17 interviews in homogeneous, narrow studies. Guest, Bunce and Johnson (2006) saw saturation within the first 12 of 60 interviews. These are study findings, not a law and not a demand test.
Is 5, 12, or 20–30 a magic number?
No. Nielsen’s 5 is usability testing of a design (~85 percent of usability problems; several small rounds), not discovery. Guest, Bunce and Johnson caution that 12 fits relatively homogeneous groups and focused aims. Griffin and Hauser’s 20–30 is a hypothesis from one product category, per segment. Adopt a starting rule or rewrite it before the first call.
Does saturation mean the idea is validated?
No. Saturation means a full batch added no new important themes. It does not mean anyone will pay, click, or stay. After the rule fires, run a named willingness-to-pay, smoke, fake-door, landing-page, design-partner, or earlyvangelist test.
What if I still hear new needs after 20–30 interviews in one segment?
The segment is probably too broad. Split it. Griffin and Hauser hypothesized 20–30 one-on-ones for 90–95 percent of needs in their category, per segment. Even the 30 customers in that study did not capture every need; their model estimated those 30 customers accounted for 89.8 percent of all needs, because low-probability needs were still missing. A new segment beats a 40th interview in the old one.
Is a Yibud score a saturation measure? Do I need an account?
No. Yibud scores come from a deterministic rule engine on your idea inputs. Saturation comes from coding need and pain phrases in one segment. You do not need an account. The method on this page stands alone. Run the free analyzer if you want a flashlight on a weak customer assumption before you write the rule. Optional language-model text only polishes prose.
Write the stopping rule before the next calendar hold
Yibud scores weak customer assumptions with a deterministic engine. You write segment, batch, saturation check, cap, and action while you are still honest — then a count is coverage, not a vibe.
Generate a Yibud report →