YibudYibudBlog indexAnalyze

Customer Research

How to Know When Customer Interviews Are Enough (A Stop-Rule for Solo Founders)

A practical stop-rule for early-stage customer interviews: when to stop, how to recognise the saturation point, why twelve is usually the right number for a single segment, and the four failure modes that keep founders interviewing long after the evidence has stopped arriving — including the difference between code saturation and meaning saturation and the named Guest, Bunce & Johnson (2006) study that produced the empirical baseline.

· Yibud· 13 min read

On this page

Quick answer

A solo founder testing a single customer segment has produced enough customer interview evidence when the new interviews stop adding new themes — a point that, for a relatively homogeneous segment, the empirical qualitative-research literature places between the tenth and twelfth interview. The Guest, Bunce & Johnson (2006) study in Field Methods — the most-cited empirical work on interview saturation to date — found that thematic saturation occurred within the first twelve interviews in a homogeneous sample, and that the basic elements of the major themes were present as early as the sixth interview. A follow-up by Hagaman & Wutich (2017) extends the rule: when the founder's research spans multiple sites or distinct sub-segments, the saturation point moves to roughly twenty to forty interviews per segment. Hennink, Kaiser & Marconi (2017) split the concept further into code saturation (the point at which no new codes are emerging — roughly nine to seventeen interviews) and meaning saturation (the point at which the relationships between codes are fully understood — roughly sixteen to twenty-four). A founder who has reached code saturation has enough to write a pattern summary; a founder who has reached meaning saturation has enough to write a confident decision.

The rest of this article is the discipline around that number: how to recognise when an interview is producing new evidence versus repeating a known theme, how the rule changes for multi-segment and B2B research, and the four failure modes — over-research, segment-switching, hypothetical-commitment, and confirmation drift — that keep founders interviewing long after the evidence has stopped arriving.

The full interview script (the questions themselves, organised by what they need to learn) is in Customer Interview Questions for Startup Validation. The discipline of behaviour-over-opinion is in The Mom Test Explained for Solo Founders. This article covers the stop rule — the part the other two do not.

Key takeaways

  • Saturation is an empirical result, not a target. Guest, Bunce & Johnson (2006) measured it: in a 60-interview study of a homogeneous population, no new themes emerged after the twelfth interview, and the basic elements of the major themes were present as early as the sixth. The number is a finding, not a slogan.
  • Code saturation and meaning saturation are different. Hennink, Kaiser & Marconi (2017) measured code saturation at 9–17 interviews and meaning saturation at 16–24 interviews. A founder writing a pattern summary needs code saturation; a founder writing a confident decision needs meaning saturation.
  • Multi-segment research pushes the number up. Hagaman & Wutich (2017) found that cross-site or cross-cultural research needs roughly 20–40 interviews per segment to surface the metathemes that cut across them. A founder researching three sub-segments should plan for 30–60 interviews, not 12.
  • The signal that you're done is the interview itself. The founder who finishes an interview and cannot name a single new theme is at saturation. The interview is research debt, not research. Rob Fitzpatrick's discipline in The Mom Test (2013) — "compliment and deflect" — assumes saturation is the normal state and new evidence is the surprise.
  • The four failure modes have names. Over-research (interviewing past saturation), segment-switching (treating each new sub-segment as a fresh start), hypothetical-commitment (counting "I would use this" as evidence), and confirmation drift (counting only the interviews that confirmed the hypothesis). Each one feels productive while it's happening. None of them produces a decision.
  • The interview loop is a means, not the work. The interview is a five-to-fifteen-hour investment whose output is a smaller to-do list, not a transcript. Steve Blank's Four Steps to the Epiphany (2005) and Eric Ries's The Lean Startup (2011) both treat the interview as a tool whose product is the assumption status it updates, not the conversation itself.

Why "how many interviews is enough" is the question every founder gets wrong

Most first-time founders treat the customer interview as the work. They schedule conversations, prepare questions, take notes, transcribe, and feel productive. The problem is that the interview is a tool, not an output. The output is the assumption status it updates — confirmed, refuted, unresolved — and the next experiment that status implies. A founder who runs twenty-five interviews without changing an assumption has produced twenty-five conversations and zero evidence.

Three things make this trap easy to fall into.

Interviews feel like progress. A calendar full of customer calls looks like a startup being built. The visual is misleading: a conversation is only valuable if it changes an assumption. A conversation that repeats a theme the founder already knew is a polite waste of both people's time.

The number "10–20 interviews" is everywhere and means nothing. It shows up in Steve Blank's slides, Y Combinator's advice pages, McKinsey decks, and most lean-startup books. The number is a rule of thumb that has been copied so often it now functions as a slogan rather than a finding. The empirical literature is sharper than that, and the sharper version is the one a founder should plan around.

The polite-yes problem makes interviews feel more productive than they are. Rob Fitzpatrick's The Mom Test (2013) is the canonical treatment. A polite audience will agree a problem is important, agree a price is reasonable, and agree they would buy the product — all in the same friendly conversation. None of that is evidence. Evidence is what the customer has already done, not what they would do in a hypothetical future. A founder who is not measuring behaviour is at risk of treating politeness as progress.

The fix is a stop rule. The founder names, before each interview, the assumption being tested and the decision that follows. After the interview, the assumption status is updated, and the decision is either made or scheduled for the next interview. When three interviews in a row leave an assumption unresolved, the next move is a different test (a landing page, a paid pilot, a concierge delivery) — not interview number fourteen.

What "saturation" actually means

Saturation is the point at which the new interviews stop producing new themes, new objections, or new information that the founder did not already have. It is a property of the interview-evidence stream, not a target number the founder should hit.

The qualitative-research literature operationalises saturation in two distinct ways, and the distinction matters.

Code saturation is the point at which no new codes are emerging — i.e., the founder has stopped hearing a new type of complaint, a new workaround, or a new segment definition. Hennink, Kaiser & Marconi (2017) found code saturation in 9–17 interviews in a 25-interview study. A founder at code saturation can write a pattern summary; they can list the recurring themes, the recurring workarounds, and the recurring objections. They cannot yet write a confident decision because the relationships between the codes are not yet fully understood.

Meaning saturation is the point at which the relationships between the codes are fully understood — i.e., the founder can describe not only the themes but also how the themes connect, who feels them most acutely, and what would change the founder's plan. Hennink, Kaiser & Marconi (2017) found meaning saturation in 16–24 interviews. A founder at meaning saturation can write a confident decision; they can name the assumption whose failure would invalidate the plan, and they can describe the segment, the channel, and the price with enough specificity to test.

The Guest, Bunce & Johnson (2006) finding — saturation within twelve interviews in a homogeneous sample — refers to thematic saturation, the point at which the major themes are fully present and the additional interviews are repeating known material. It is closer to code saturation than to meaning saturation, which is why a founder using the "twelve-interview rule" without distinguishing the two kinds of saturation may reach code saturation without reaching meaning saturation.

The practical implication is that the founder's target depends on the decision being made. A founder who needs a pattern summary to write a landing page can stop at code saturation (roughly twelve interviews in a single segment). A founder who needs to write a confident go/no-go decision should plan for meaning saturation (roughly sixteen to twenty-four interviews) — or, if the segment is complex, treat twelve as the end of the first pass and use the next four to twelve interviews to test the relationships between the themes.

The four signs you have reached saturation

The number is the easy answer. The harder answer is the recognition: how does the founder know, sitting in front of an interviewee, that this conversation is not going to produce new evidence?

The four signals below each have a corresponding behaviour. A founder who notices any one of them has probably reached saturation and is incurring research debt by continuing.

1. The conversation is repeating a known theme

The most direct signal. The interviewer's notes after the interview contain the same three complaints, workarounds, and price ranges the previous three interviews also contained. The customer is not a poor interviewee; the founder has extracted everything this segment is producing.

The fix is to stop. Either the assumption the interview was designed to test is now well-supported, and the next move is a different test (a landing page, a paid pilot), or the assumption is unsupported and the segment is wrong.

2. The new interviews are not changing the founder's segment definition

A second signal. If interview three narrowed the segment to "independent marketing consultants with retainer clients in the US," and interviews four through seven keep confirming that definition without refining it, the founder is at saturation for that segment. The remaining uncertainty is no longer about whether the segment has the problem; it is about whether the founder can reach them at a cost the model supports — and that is a channel test, not another interview.

3. The customer is volunteering information the founder has not yet recorded

A subtle signal. The interview is moving faster than the founder can take notes because the customer is volunteering things — workarounds, costs, frustrations, recent events — that the founder did not have to ask about. This is the sign of an under-saturated interview: there is still a lot to learn, and the conversation is rich.

The opposite signal — the founder is asking every question on the script and the customer is offering short answers — is the sign of saturation. The customer has said what they have to say. The interview is now extracting politeness, not evidence.

4. The founder is interviewing outside the named segment

The final signal, and the most common one. A founder who has run twelve interviews in the named segment and cannot find a thirteenth person who fits the segment definition is a founder who has exhausted the segment. Two temptations follow: broadening the segment ("any consultant, anywhere"), or moving to a different segment ("what about small agencies?"). Both feel like progress; both produce a different kind of evidence, and neither is comparable to the first twelve interviews. The right move is to stop, write the pattern summary from the first twelve, and decide whether to switch segments deliberately (with a new saturation target) or to leave the segment as is and move on.

The empirical baseline — what the named studies actually found

The interview-saturation numbers founders use come from a small set of empirical qualitative-research papers. Citing them is the difference between a number and a finding.

Guest, Bunce & Johnson (2006), "How Many Interviews Are Enough? An Experiment with Data Saturation and Variability," Field Methods, 18(1), 59–82. The authors analysed 60 in-depth interviews on sexual behaviour and risk with women in two West African countries. They tracked the appearance of new thematic codes across successive interviews and found that no new themes emerged after the twelfth interview, and the basic elements of the major themes were present as early as the sixth. The study is one of the most-cited qualitative methodology papers (over 3,900 citations as of recent counts) and is often invoked to justify small qualitative samples. Two caveats apply: the population was relatively homogeneous, and the researchers were experienced qualitative researchers familiar with the context. A founder doing customer discovery in a heterogeneous customer base should treat the twelve-interview number as a floor, not a ceiling.

Hennink, Kaiser & Marconi (2017), "Code Saturation Versus Meaning Saturation: How Many Interviews Are Enough for Qualitative Research?" Qualitative Health Research, 27(4), 591–608. The authors studied a 25-interview dataset on international reproductive health and tracked two separate saturation points: code saturation (the point at which no new codes are emerging) and meaning saturation (the point at which the relationships between codes are fully understood). Code saturation occurred at roughly nine interviews; meaning saturation occurred at roughly sixteen to seventeen interviews. The distinction is the practical contribution of the paper. A founder who needs to write a pattern summary can stop at code saturation; a founder who needs to write a confident decision should plan for meaning saturation.

Hagaman & Wutich (2017), "How Many Interviews Are Enough to Identify Metathemes in Multisited and Cross-cultural Research? Another Perspective on Guest, Bunce, and Johnson's (2006) Landmark Study," Field Methods, 29(1), 23–41. The authors extended the Guest et al. methodology to a cross-cultural study of water issues in four sites (Bolivia, New Zealand, Fiji, and the US) with 132 respondents. They found that within-site saturation followed the Guest et al. pattern (sixteen interviews per site for homogeneous sub-populations), but cross-site metathemes — themes that cut across sites — required twenty to forty interviews per site to surface. The implication for a founder researching multiple sub-segments is direct: the saturation target scales with the heterogeneity of the research question.

Together, the three studies give the founder a planning rule rather than a slogan. For a single-segment, single-question founder, twelve interviews is a defensible baseline; sixteen to twenty-four produces meaning saturation. For a multi-segment founder, the target is twenty to forty per segment, with a total of thirty to sixty across the research project. Numbers above that are usually research debt.

Segment-specific stop rules

The numbers above are baselines; the right number depends on the segment the founder is researching.

Single-segment, single-question (the simplest case)

Plan for twelve interviews. Stop when three interviews in a row do not produce a new theme. The output is a pattern summary naming the recurring themes, the segment definition, and the unresolved assumptions. This is the typical indie-hacker case.

Single-segment, multi-question

Plan for sixteen to twenty-four interviews. The founder is testing multiple assumptions about the same segment (frequency of the problem, severity, willingness to pay, willingness to switch from the current workaround). Code saturation may still occur at twelve, but meaning saturation — understanding how the themes relate to each other — requires the second pass. The output is a confident decision naming the strongest assumption, the weakest assumption, and the next experiment.

Multi-segment, single-question

Plan for twenty to forty interviews per segment, with a minimum of three segments. Total: sixty to one hundred and twenty interviews. This is the typical B2B or vertical SaaS case where the founder needs to know whether the problem exists across multiple buyer roles or industries. The output is a segment map naming which sub-segments have the problem, which do not, and where the highest-value segment sits.

Multi-segment, multi-question

Plan for the multi-segment rule above, plus the meaning-saturation extension. Total: sixty to one hundred and fifty interviews across all segments. This is the typical consulting-firm or agency case and is rarely the right level of effort for a solo founder. The advice for solo founders: pick the single most promising segment and the single most important question, and run the simpler rule first.

B2B with hard-to-reach buyers

The rule is the same in principle, but the channel changes. Cold email and warm intros replace community recruiting. The bar for "an interview" is higher — a thirty-minute call with a director-level buyer is harder to schedule than a fifteen-minute call with a consumer — so the founder should plan for a longer calendar even if the saturation point is reached at a smaller total number. The full discipline is in How to Validate a B2B Startup Idea Before You Build It.

Three labeled hypothetical examples

The examples below are illustrative. The segments, the numbers, and the decisions are constructed to show what saturation looks like in practice; they are not descriptions of any specific real company. Real founders should adapt the rule to their own segment, hypothesis, and decision.

Example 1 — Single-segment B2B SaaS (hypothetical)

  • Idea. A SaaS tool that automates client-report generation for independent marketing consultants in the US with 3–15 retainer clients.
  • Interviews.
    • Interviews 1–4: four different consultants describe the same problem unprompted (manual reporting, 4–6 hours per month), name the same workaround (spreadsheet + template), and quote the same cost ($200–400 in unbillable hours per month).
    • Interviews 5–7: three more consultants confirm the same pattern. No new themes. One mentions a tool they tried and abandoned; the abandonment reason is new.
    • Interviews 8–9: two more consultants. The pattern is now firmly established; the only new information is a specific price point ($149/month) that two customers say is below what they would pay.
    • Interview 10: a consultant who fits the segment but does not have the problem (the agency is too small, the founder does the reports herself).
  • Saturation signal. After interview 9, no new themes are emerging. Interview 10 refines the segment but does not add a new theme.
  • Stop decision. The founder stops at interview 10. Code saturation is reached. The pattern summary is written; the next experiment is a landing page test with the $149 price point. Meaning saturation is not yet reached and will be re-tested after the landing page produces a signal.

Example 2 — Multi-segment B2C consumer app (hypothetical)

  • Idea. A daily mindfulness app for adults who have completed a paid meditation course.
  • Interviews.
    • Interviews 1–6: six conversations with members of the segment. Three describe wanting a daily practice but struggling to maintain it; three describe being satisfied with their current course-based routine. A pattern is forming.
    • Interviews 7–12: six more conversations. The pattern holds. Two customers mention a specific feature they would want (a five-minute morning version); no other new themes.
    • Interviews 13–18: six more conversations with the same segment. No new themes. Customers repeat the same complaints and the same workarounds.
    • Sub-segment check: the founder interviews six former course graduates who completed their course more than a year ago. The pattern reverses: this sub-segment has moved on and is no longer interested in a daily mindfulness practice.
  • Saturation signal. Interviews 13–18 produced no new themes in the original segment; the sub-segment check produced a clean negative signal in the second segment.
  • Stop decision. The founder stops at interview 18 in the original segment (meaning saturation reached for the question "do former course graduates want a daily practice?") and at interview 6 in the sub-segment. The pattern is clear: the product has a six-month window after course completion and no demand beyond it. The next experiment is a landing page targeted at recent course graduates, not a broader audience.

Example 3 — Hard-to-reach B2B buyer (hypothetical)

  • Idea. An enterprise SaaS tool that automates compliance reporting for mid-market banks in the EU.
  • Interviews.
    • Interviews 1–4: four chief compliance officers at banks with €1–10 billion in assets. All four describe the same problem (manual reporting, multi-week process), name the same workaround (a mix of spreadsheets and a large consultancy), and quote the same cost (€50–150K per year in consultancy fees). A pattern is forming.
    • Interviews 5–8: four more CCOs. The pattern holds; one new theme emerges (a specific regulatory change in 2024 that has increased the workload).
    • Interviews 9–12: four more CCOs. No new themes. The 2024 regulatory change is the most-cited reason for the problem; the workaround pattern is uniform.
  • Saturation signal. After interview 12, no new themes are emerging. The pattern is robust across twelve interviews with hard-to-reach buyers; the cost of scheduling each interview is high (an average of three weeks of outreach per conversation).
  • Stop decision. The founder stops at interview 12. Code saturation is reached; meaning saturation is reached for the problem question, and the next experiment is a paid pilot with one or two of the interviewees to test the price point and the workflow. Total interview investment: roughly 36 weeks of outreach scheduling compressed into a four-month calendar. The decision is to move to a paid pilot, not to interview more buyers.

Comparison table — what each interview actually produces

Interview #Typical outputWhat it cannot tell youStop-or-continue signal
1–3First themes; first objections; first workaroundsWhether the themes will repeatContinue — sample is too small to confirm anything
4–6Themes begin to repeat; basic elements of major themes present (Guest et al. 2006)Whether the relationships between themes will holdContinue unless 4–6 all produced the same themes with no new material
7–9Patterns are clear; first sub-segment signals appearWhether the sub-segments are real or noiseContinue unless 7–9 produced no new themes
10–12Code saturation is plausible; pattern summary can be writtenWhether meaning saturation is reachedStop the first pass; re-test after a landing page or pilot produces a signal
13–17Meaning saturation is plausible; relationships between themes are understoodWhether the pattern holds outside the segmentStop and write the decision unless a new sub-segment is being researched
18–24Meaning saturation is reached; confident decision can be writtenWhether the segment is large enough to support a businessStop and move to a pricing or channel test
25+Repetition; confirmation drift; segment-switchingNothing the previous interviews did not already produceStop. Further interviews are research debt, not research

The right column is the discipline. The founder should be able to look at the interview log and tell, for each interview, what it produced that the previous one did not. An interview that produces nothing new is the saturation signal.

Common mistakes

The four failure modes below account for most of the over-research that keeps founders interviewing past saturation. Each one feels productive while it's happening. None of them produces a decision.

  1. Over-research. The founder schedules interview after interview because stopping feels like giving up. The customer-interview log is filling up; the founder's to-do list is not. Fix: write the pattern summary at interview 10 even if the founder is not sure it is right. The summary will be wrong; the next experiment will reveal how. A wrong pattern summary is more useful than a perfect transcript.

  2. Segment-switching. The founder has saturated the named segment and cannot find a thirteenth interviewee who fits. The temptation is to broaden the segment ("any consultant, anywhere") or to switch segments ("what about small agencies?"). Both feel like progress; both produce a different kind of evidence that is not comparable to the first twelve interviews. Fix: when the named segment is exhausted, the founder makes a deliberate decision to switch segments (with a new saturation target) or to leave the segment as is. Segment-switching by drift is research debt.

  3. Hypothetical-commitment. The founder is counting "I would use this" or "I would pay for this" as evidence. The customer is being polite. Fix: the next interview replaces hypothetical questions with behaviour questions. "Tell me about the last time you had this problem" replaces "Would you use a tool that fixed this." The customer's memory is evidence; their hypothetical opinion is not. The full discipline is in The Mom Test Explained for Solo Founders.

  4. Confirmation drift. The founder counts the interviews that confirmed the hypothesis and ignores the interviews that did not. The interview log grows; the pattern summary becomes one-sided. Fix: write the pattern summary before the next interview. The summary must include both confirming and disconfirming evidence. A summary that reads like a sales pitch is a summary that has been written by confirmation drift.

A fifth, subtler failure mode is interview inflation — treating each interview as a milestone because it is the eleventh or the fifteenth or the twentieth. The number is not a target; it is a finding. A founder who has reached saturation at interview 8 is not "behind"; they are ahead. A founder who has not reached saturation at interview 14 is not "ahead"; they have a more complex research question than they thought.

Limits of the stop rule

The empirical literature is a baseline, not a guarantee. Three limits apply.

Homogeneous samples saturate faster than heterogeneous samples. The Guest et al. (2006) twelve-interview finding is for a relatively homogeneous population. A founder researching a heterogeneous customer base — multiple industries, multiple use cases, multiple geographies — should expect saturation later, in the Hagaman & Wutich (2017) twenty-to-forty range. The right move is to narrow the segment before counting interviews, not after.

Trained researchers saturate faster than first-time founders. The Guest et al. (2006) study was conducted by experienced qualitative researchers familiar with the context. A first-time founder is learning the discipline while doing the research. The implication is that a first-time founder should plan for the upper end of the saturation range, not the lower end. Twelve interviews is a floor for a first-time founder in a single segment, not a ceiling.

Saturation is a property of the question, not the population. A founder researching a narrow question ("do independent marketing consultants in the US have a manual-reporting problem?") saturates faster than a founder researching a broad question ("what is the workflow of independent marketing consultants in the US?"). The discipline is to narrow the question before counting interviews. A founder who has spent twelve interviews on a broad question has produced a context map, not a pattern summary.

The interview loop is a means, not the work. Teresa Torres's Continuous Discovery Habits (2021) argues that customer research is a weekly practice, not a one-time gate. The stop rule is for the pre-build validation phase; after the build, customer research becomes part of the weekly cadence. The full loop is in Continuous Discovery.

FAQ

How many customer interviews should a startup conduct?

For a solo founder researching a single segment with a single question, twelve interviews is a defensible baseline (Guest, Bunce & Johnson 2006); sixteen to twenty-four interviews produces meaning saturation (Hennink, Kaiser & Marconi 2017). For a multi-segment founder, plan for twenty to forty interviews per segment (Hagaman & Wutich 2017). The right number depends on the segment's heterogeneity, the question's breadth, and the founder's experience.

How do I know when to stop interviewing customers?

The signal is the interview itself. If the founder finishes an interview and cannot name a single new theme, objection, or workaround that the previous three interviews did not also produce, the founder is at saturation. The next interview will repeat known material. The right move is to write the pattern summary from the interviews already run and design the next experiment (a landing page, a paid pilot, a concierge delivery) to test the strongest unresolved assumption.

What is data saturation in qualitative research?

Data saturation is the point at which additional interviews stop producing new themes, codes, or relationships between codes. The concept was operationalised empirically by Guest, Bunce & Johnson (2006) in Field Methods, who found that thematic saturation occurred within twelve interviews in a homogeneous sample. The follow-up literature distinguishes code saturation (no new codes; 9–17 interviews per Hennink, Kaiser & Marconi 2017) from meaning saturation (no new relationships; 16–24 interviews).

Is ten customer interviews enough?

For a single-segment founder with a narrow question, ten interviews is at or near the lower bound of the empirical saturation range. It is enough to write a pattern summary. It is not enough to write a confident go/no-go decision; the founder should plan for a second pass of six to fourteen more interviews, or move to a different test (landing page, paid pilot) that can produce a stronger signal.

Is twenty customer interviews too many?

For a single-segment founder with a narrow question, twenty interviews is past the code-saturation point and at or past the meaning-saturation point. The twenty-first interview in the same segment is almost certainly research debt. The exceptions are multi-segment research (Hagaman & Wutich 2017 — twenty to forty interviews per segment) and complex heterogeneous populations (where meaning saturation can extend to twenty-four or beyond per Hennink, Kaiser & Marconi 2017).

What if my interviews keep producing new themes?

Two possibilities. Either the founder has a more heterogeneous customer base than they realised and should plan for a longer saturation run (twenty to forty per segment), or the founder's question is too broad and should be narrowed. The fix is rarely "do more interviews"; it is usually "narrow the segment or narrow the question, then continue."

How do I run customer interviews without leading the witness?

Replace hypothetical questions with behaviour questions. "Would you use this?" produces a polite yes. "Tell me about the last time you had this problem" produces a memory, and the memory is the evidence. The full discipline, including the twelve-question script and the failure modes, is in The Mom Test Explained for Solo Founders and Customer Interview Questions for Startup Validation.

What should I do after customer interviews are saturated?

Convert each interview's findings into an assumption status — confirmed, refuted, or unresolved. Then run a pattern check across all the interviews: which assumptions are supported by more than one customer, which are still unresolved, and which have been refuted. The unresolved assumptions are the next experiments. The refuted assumptions are the changes the founder should make before building. The supported assumptions are the assumptions the founder can act on with reasonable confidence. Then connect the assumptions to the broader validation framework — Lean Startup Validation covers the experiment loop, and How to Validate a Startup Idea Before Building is the broader framework this discipline sits inside.

Summary

Knowing when customer interviews are enough is the discipline of recognising saturation: the point at which additional interviews stop producing new evidence. The empirical qualitative-research literature places this point at twelve interviews in a homogeneous single-segment sample (Guest, Bunce & Johnson 2006), sixteen to twenty-four for meaning saturation in the same sample (Hennink, Kaiser & Marconi 2017), and twenty to forty per segment in multi-segment research (Hagaman & Wutich 2017). The number is a finding, not a target — a founder who has reached saturation at interview 8 is ahead, not behind.

The signal is the interview itself. A founder who finishes an interview and cannot name a single new theme is at saturation. The four failure modes — over-research, segment-switching, hypothetical-commitment, and confirmation drift — keep founders interviewing past the point where the evidence has stopped arriving. The interview loop is a means, not the work; the output is the assumption status it updates, not the conversation itself.

A founder who has run twelve interviews in a single segment has enough to write a pattern summary. A founder who has run sixteen to twenty-four has enough to write a confident decision. A founder who has run twenty-five in the same segment has produced research debt. The next move after saturation is a different test — a landing page, a paid pilot, a concierge delivery — not interview number twenty-six.

Sources

  • Guest, G., Bunce, A., & Johnson, L. (2006). How Many Interviews Are Enough? An Experiment with Data Saturation and Variability. Field Methods, 18(1), 59–82. DOI: 10.1177/1525822X05279903. The empirical baseline: thematic saturation occurred within the first twelve interviews in a homogeneous sample of 60 interviews; basic elements of major themes were present as early as interview 6.
  • Hennink, M. M., Kaiser, B. N., & Marconi, V. C. (2017). Code Saturation Versus Meaning Saturation: How Many Interviews Are Enough for Qualitative Research? Qualitative Health Research, 27(4), 591–608. DOI: 10.1177/1049732316665344. Distinguishes code saturation (9–17 interviews) from meaning saturation (16–24 interviews); a founder writing a pattern summary can stop at code saturation, a founder writing a confident decision should plan for meaning saturation.
  • Hagaman, A. K., & Wutich, A. (2017). How Many Interviews Are Enough to Identify Metathemes in Multisited and Cross-cultural Research? Another Perspective on Guest, Bunce, and Johnson's (2006) Landmark Study. Field Methods, 29(1), 23–41. DOI: 10.1177/1525822X16640447. Extends the Guest et al. methodology to multi-site research: 20–40 interviews per site for cross-site metathemes.
  • Fitzpatrick, R. (2013). The Mom Test: How to Talk to Customers & Learn If Your Business is a Good Idea When Everyone is Lying to You. The polite-yes problem and the behaviour-over-opinion discipline that anchors the customer interview loop. Summary at momtestbook.com.
  • Blank, S. (2005). The Four Steps to the Epiphany: Successful Strategies for Products that Win. Customer Development methodology; the interview loop as the entry condition for the build. Summary at steveblank.com.
  • Ries, E. (2011). The Lean Startup: How Today's Entrepreneurs Use Continuous Innovation to Create Radically Successful Businesses. Build-Measure-Learn loop; the interview as the input to the assumption test. Reference at theleanstartup.com.
  • Torres, T. (2021). Continuous Discovery Habits: Discover What Your Customers Need Before They Do. Customer research as a weekly practice, not a one-time gate. Reference at continuousdiscoveryhabits.com.

Next action

If the founder has not yet run a customer interview, the next concrete action is to schedule one this week. A single thirty-minute conversation with one member of the named customer segment, asking about their last encounter with the problem the founder's idea is meant to solve. Not a survey, not a focus group, not a Slack message. A real conversation, with notes taken during rather than after.

If the founder has already run interviews, the next concrete action is to write the pattern summary from the interviews they have. The summary should name the recurring themes, the segment definition, the recurring workarounds, and the unresolved assumptions. If the founder cannot fill in any one of those four fields from the interviews they have run, the next interview is the one that closes the gap. If the founder can fill in all four, the next move is a different test — a landing page, a paid pilot, a concierge delivery — not another interview.

Test your own idea

Describe your idea, answer five short questions, and get a structured 8-dimension report — free, no signup.

Continue learning

Where to go from here

These pieces are grouped by topic, not publication date — pick the one that matches the question you are working on right now.

See every article on startup validation in one place.

Open the Startup Validation hub →