YibudYibudBlog indexAnalyze

Validation Guide

Lean Startup Validation: How To Choose, Design and Interpret Validation Experiments

How to use Eric Ries's Build-Measure-Learn loop to pick the next validation experiment — including the five-question decision framework, decision-threshold thinking, three labeled hypothetical examples, and the failure modes that distort the loop before the founder notices them.

· Updated · Yibud· 17 min read

On this page

Quick answer

Lean Startup validation is the discipline of testing startup assumptions through short, iterative experiments before committing to a full build. Eric Ries formalized the method in The Lean Startup (2011) as a loop — Build → Measure → Learn — that turns each assumption into a hypothesis, each hypothesis into the smallest experiment that can produce a signal, and each signal into a one-sentence learning. The output is validated learning, not a launch.

The hardest part is not running a single experiment. It is choosing the next one. The framework below picks the next experiment by answering five questions about the assumption you are testing, the uncertainty around it, the behavior that would change your mind, the cheapest way to observe it, and the result that would force a new decision.

The smallest "build" can be a landing page, a customer interview, a concierge MVP, a Wizard-of-Oz service, a prototype, or a paid pilot — whichever is the cheapest credible artifact for the assumption being tested. The right method matches the assumption. Ries's definition of the MVP — "a version of a new product that allows a team to collect the maximum amount of validated learning about customers with the least effort" — is the same principle at product scale.

Key takeaways

  • Lean Startup is a learning method, not a launch method. A test whose result cannot come back negative is a launch event in disguise; it cannot produce validated learning.
  • The first thing to validate is the riskiest assumption. Steve Blank's Four Steps to the Epiphany (2005) and the Customer Development methodology introduced the riskiest-assumption test (RAT): rank each assumption by uncertainty × impact, then test the one whose failure would invalidate the rest of the plan.
  • "Build" does not mean "build a product." A landing page, an interview, a concierge service, a Wizard-of-Oz setup, a prototype, or a paid pilot are all valid Lean Startup "builds." The smallest credible artifact for the assumption is the right one.
  • Behavior, not opinion, is the only honest signal. Rob Fitzpatrick's The Mom Test (2013) makes the polite-yes problem canonical: customers say yes to be polite and then do nothing. The loop measures what the customer does, not what the customer says.
  • Validated learning is the unit of progress. The sentence you write after each loop must be empirical (grounded in behavior), falsifiable (could be contradicted by future evidence), and actionable (points to the next test). A founder who cannot write that sentence has not run an experiment.
  • Lean Startup is not a verdict on success. It is a method for narrowing the range of things the founder could still be wrong about. A clean experiment produces a smaller to-do list, not a green light.

What Lean Startup validation is — and is not

Lean Startup validation is a methodology for learning through iterative experiments. It is one way to run startup validation; it is not the only one, and it does not cover the full validation discipline.

The relationship matters for the rest of this page:

  • Startup validation is the broader discipline of testing whether an idea has enough evidence to justify continued investment. It includes problem validation, customer discovery, willingness-to-pay testing, MVP validation, and product-market-fit testing. The canonical framework is How to Validate a Startup Idea Before Building.
  • Lean Startup is the experimentation methodology. It operates inside startup validation; it does not replace it. The vocabulary is defined in the Startup Validation glossary.
  • Startup MRI is a structured second opinion that surfaces the riskiest assumption for one specific idea across nine dimensions. It does not run experiments; it tells you which assumption the loop should test first.

This page's purpose is narrower than those three. It covers how to choose, design, and interpret Lean Startup experiments. The other pages in the cluster cover the related rungs.

The riskiest-assumption test

A founder typically has between ten and thirty assumptions stacked behind an idea. The most common Lean Startup mistake is testing them in the wrong order. The right order is not "the cheapest first" — it is the riskiest first, where risk equals uncertainty multiplied by impact.

A practical priority order, in roughly the sequence most solo founders should consider:

  1. Problem — Does the customer experience this problem, and does it recur often enough to act on?
  2. Customer — Is the named segment specific enough to find ten of them this week?
  3. Urgency — Is the problem acute enough that the customer is already solving it badly?
  4. Existing alternatives — What is the customer doing today, and at what cost?
  5. Willingness to pay — Will the customer commit real money at a price that supports the model? See Willingness to Pay Validation.
  6. Acquisition / distribution — Can the customer be reached through a named channel at a cost the model supports? See Distribution Channels Ranked for Solo Founders.
  7. Retention — Does the customer come back without prompting, at the product's natural frequency? See Product-Market Fit Validation.
  8. Solution — Does the proposed solution actually address the problem? See Problem-Solution Fit Validation.

The order is not universal. A founder who already has strong evidence on rungs one through three may validly start at rung five. A founder in a regulated category may need to validate the regulatory pathway before anything else. The principle is constant: test the assumption whose failure would invalidate the rest of the plan, not the assumption that is easiest to test.

How to choose your next validation experiment

Once the riskiest assumption is identified, the next decision is which experiment to run. The framework below picks the experiment by answering five questions in order. The questions turn an assumption into an experiment design.

The five questions

  1. What assumption could make the idea fail? A single sentence in the form "We believe [customer] will [behavior] when [condition]." If the assumption cannot be written, it is not yet ready to test.
  2. How uncertain is that assumption? Score from 1 (very confident) to 5 (no evidence either way). The score is honest about what you currently believe, not about what you hope.
  3. What observable behavior would reduce uncertainty? A specific action the customer takes — signing up, paying, returning, referring — that is countable and time-bounded.
  4. What is the cheapest credible experiment that could produce that behavior? The cheapest artifact that could realistically cause the named customer, in the named channel, to perform the named action. A landing page, an interview script, a concierge flow, a Wizard-of-Oz setup, a prototype, a paid pilot, or a small feature-scope MVP. Cheapest does not mean lowest absolute cost; it means lowest cost relative to the strength of the signal you need. Six build methods ranked by signal strength are in MVP Validation.
  5. What result would change the next decision? Pre-commit to what counts as useful evidence, what would be inconclusive, what would justify another experiment, and what would force a reconsideration. The decision rule is written before the experiment runs.

The output of the five questions is a one-page experiment design. The same artifact is reusable as a template for the next experiment; only the assumption and the named customer change.

Decision thresholds

Before running an experiment, define four outcomes. The thresholds are not universal numbers; they depend on the product, the customer, the cost of the experiment, and the decision being made. What is "good enough" for a $200 landing-page test is not good enough for a six-month engineering build.

For each experiment, write down:

  • Useful evidence. The result that would actually change the next decision. A landing page test that converts at 8% on 200 qualified visitors is useful evidence. 8% on 20 unqualified visitors is not.
  • Inconclusive result. The result that would justify another experiment rather than a decision. A 2% conversion rate on 50 visitors is inconclusive; a 2% rate on 5,000 visitors is more likely real.
  • Re-run trigger. The result that would mean the experiment itself was wrong — wrong channel, wrong customer, wrong offer — and should be re-run with a changed variable.
  • Reconsider trigger. The result that would mean the assumption itself is unlikely to hold and the loop should pivot or stop. Three customers in a row saying the same negative thing is a reconsider trigger; one polite "no" is not.

A founder who skips the decision thresholds runs experiments whose results cannot come back negative. That is not validation; it is confirmation with extra steps.

Decision framework: Assumption → Experiment → Evidence → Learning → Decision

The loop runs in five sequential steps. Each step produces an artifact that can be reviewed before the next step runs.

Assumption
   ↓
Risk (uncertainty × impact)
   ↓
Experiment (cheapest credible build)
   ↓
Evidence (observable customer behavior)
   ↓
Learning (one empirical, falsifiable, actionable sentence)
   ↓
Decision (continue, change assumption, change customer,
          change solution, change channel, stop)

The six named decisions are the full set of legitimate outputs of a clean experiment. Stop is a legitimate outcome; "the experiment did not work" is the most useful kind of evidence the loop can produce.

Three labeled hypothetical examples

The examples below are illustrative. The numbers and customer segments are constructed to show what the framework looks like in practice; they are not descriptions of any specific real company. Real founders should adapt the assumption, the named customer, the experiment, and the decision rule to their own context.

Example 1 — B2B SaaS (hypothetical)

  • Assumption. "We believe property managers with 10–50 units will pay $99/month for a software tool that automates their monthly reporting."
  • Risk. High. No prior evidence of payment intent; the customer is reachable but the price point is unproven.
  • Cheapest credible test. Five problem interviews with named property managers, followed by a two-week concierge MVP using a shared spreadsheet and an email parser, delivered manually to up to three customers at the planned price.
  • Evidence to observe. Unprompted problem stories in the interviews (count, severity, frequency, current workaround cost). In the concierge: number of customers who pay the first invoice at the end of the month, number who actively use the workflow past day seven.
  • Decision rule.
    • Useful evidence: 3 of 5 interviewees describe the problem unprompted AND 2 of 3 concierge customers pay the first invoice.
    • Inconclusive: Mixed interview signals, or 1 of 3 concierge customers pays.
    • Re-run trigger: Interviewees describe the problem but none agree to the concierge — likely a pricing or trust issue, not a problem issue.
    • Reconsider trigger: 0 of 5 interviewees describe the problem unprompted, or 0 of 3 concierge customers pay after two weeks of free use.
  • Next decision after the test. If useful evidence: continue with a narrower assumption about which segment of property managers is most engaged, and design the next experiment to test retention. If inconclusive: change the customer segment or change the offer. If reconsider: stop or change the problem.

Example 2 — Consumer product (hypothetical)

  • Assumption. "We believe adults who have completed a paid meditation course will use a daily mindfulness app at least four days per week for at least eight weeks."
  • Risk. Medium-high. The customer is identified; the retention frequency is unproven.
  • Cheapest credible test. A two-week prototype test with fifteen recruited users from the named segment, with a retention dashboard that tracks daily opens, completed sessions, and day-7 return.
  • Evidence to observe. Day-1 retention, day-7 retention, day-14 retention. Frequency of completed sessions per active user.
  • Decision rule.
    • Useful evidence: 8 of 15 users complete onboarding AND 4 of those return on day 7 AND 3 of those return on day 14.
    • Inconclusive: 5–7 of 15 complete onboarding — the signal is directionally right but not strong enough for a multi-month build.
    • Re-run trigger: Low onboarding completion but high day-7 retention among completers — the onboarding itself may be the problem, not the retention assumption.
    • Reconsider trigger: Fewer than 3 of 15 complete onboarding, or 0 of those return on day 7.
  • Next decision after the test. If useful evidence: change the assumption to a narrower segment (for example, "people who finished a specific paid course in the last 90 days") and re-test. If inconclusive: refine the onboarding or test a different segment. If reconsider: pivot to a different use case or stop.

Example 3 — Marketplace (hypothetical)

  • Assumption. "We believe independent writers and small-business clients can both be acquired through the same initial paid social channel."
  • Risk. High. The single-channel assumption is the load-bearing piece; if wrong, both sides must be acquired separately.
  • Cheapest credible test. A six-week manual matching pilot. The founder runs one ad campaign aimed at both sides, takes intake from both, matches them by hand, processes three transactions per week, and records the funnel at each step.
  • Evidence to observe. Cost per qualified intake on each side. Quality of writer applications (do the best writers apply?). Number of completed transactions per week. Number of failed transactions and the reason for each failure.
  • Decision rule.
    • Useful evidence: Cost per qualified intake on both sides within 2× of each other AND 6+ completed transactions in six weeks AND fewer than 2 failed transactions.
    • Inconclusive: Large cost-per-intake gap between sides but transactions do complete — the channel works for one side but not the other.
    • Re-run trigger: Intake quality on one side is poor (low-intent applicants) — the channel reaches the wrong audience for that side.
    • Reconsider trigger: Fewer than 3 completed transactions in six weeks, or 4+ failed transactions.
  • Next decision after the test. If useful evidence: continue, run a second pilot with the named single channel, and measure retention of completed matches. If inconclusive: change the channel for the underperforming side. If reconsider: stop or change the customer definition for one side.

Worked comparison table

The table below puts the three examples side by side. Use it as a template: replace the rows with your own assumption, customer, experiment, evidence, and decision rule.

AssumptionRisk if wrongCheapest credible testEvidence to observePossible next decision
Property managers with 10–50 units will pay $99/month for automated monthly reporting (B2B SaaS)High: payment intent unproven5 problem interviews + 2-week concierge MVP at planned priceUnprompted problem stories; first-invoice payment rateContinue with narrower segment; change segment; reconsider problem
Adults who completed a paid meditation course will use a daily mindfulness app 4×/week for 8 weeks (Consumer)Medium-high: retention frequency unproven2-week prototype with retention dashboardDay-1, day-7, day-14 retention; completed sessionsContinue with narrower segment; refine onboarding; pivot or stop
Writers and small-business clients can both be acquired through one paid social channel (Marketplace)High: single-channel assumption is load-bearing6-week manual matching pilot with one campaign aimed at both sidesCost per qualified intake per side; completed vs failed transactionsContinue; change channel for one side; stop or redefine customer

The table is the artifact the loop produces. The same five columns apply to any assumption; only the entries change.

Common mistakes

Eight failure modes show up before the founder notices them. Each has a recognizable signature; each has a fix that starts with the loop rung that broke.

  • Testing the solution before the problem. The founder starts at rung eight when rungs one through four are still unverified. Write down the eight assumptions in priority order; start at the top.
  • Asking hypothetical questions instead of observing behavior. "Would you pay $99/month?" is hypothetical; "put a deposit on the next slot" is real. If a question cannot be converted to a commitment, it is the wrong question.
  • Treating compliments as demand. Beta users say nice things; polite respondents agree the problem is real. Both are enthusiasm, not evidence. Measure commitment, use, and referral behavior instead of opinion.
  • Choosing metrics after seeing the results. The metric and the threshold belong in step four of the five-question framework, before the experiment runs. A metric chosen after the data is not a result; it is a story.
  • Running experiments without a decision rule. The four outcomes (useful / inconclusive / re-run / reconsider) belong in the experiment design, not after the data arrives. A test without a decision rule is not an experiment.
  • Building too much before learning anything. The MVP becomes the test, and the founder cannot distinguish the assumption being tested from the features bundled around it. Start at the smallest build that can produce a signal.
  • Confusing activity with learning. Shipping five features in a month is activity. Writing one sentence about what was learned is learning. A founder who cannot produce the sentence has not been running experiments.
  • Ignoring negative evidence. Three customers in a row say the same negative thing; the founder updates the marketing copy and re-runs. The "no" is the most informative response in the dataset. Treat it that way.

Each of these is a rung in the loop that broke. The fix is the same in every case: go back to the rung that produced the gap, repair it, and re-run.

FAQ

What is Lean Startup validation?

Lean Startup validation is the discipline of testing startup assumptions through iterative Build-Measure-Learn experiments before committing to a full build. The method was formalized by Eric Ries in The Lean Startup (2011) and is one of several ways to run startup validation. The goal is to reduce uncertainty, not to predict success.

How do I pick which validation experiment to run next?

Rank your assumptions by risk (uncertainty × impact) and test the riskiest first. Then answer the five questions in the framework above: what is the assumption, how uncertain are you, what behavior would change your mind, what is the cheapest credible experiment that could produce that behavior, and what result would force a new decision. The output is a one-page experiment design with a pre-committed decision rule.

What is the Build-Measure-Learn loop?

Build the smallest artifact that can produce a signal about the riskiest assumption. Measure observable customer behavior, not opinions. Learn by converting the measurement into a single falsifiable sentence about what is now known. Repeat. The loop is iterative on purpose; the first experiment rarely settles the question.

What is the difference between Lean Startup and startup validation?

Startup validation is the broader discipline of testing whether an idea has enough evidence to justify continued investment. Lean Startup is the experimentation methodology that operates inside startup validation. Startup validation includes Lean Startup as one method among several. The canonical framework that puts the loop in context is How to Validate a Startup Idea Before Building.

How do I decide when an experiment result is "good enough"?

Before running the experiment, define four outcomes: useful evidence, inconclusive, re-run trigger, and reconsider trigger. The thresholds depend on the product, the customer, the cost of the experiment, and the decision being made. A founder who skips the pre-commit cannot read the result honestly when it arrives.

What is the difference between an MVP and a validation experiment?

Ries's definition of an MVP — "a version of a new product that allows a team to collect the maximum amount of validated learning about customers with the least effort" — is a specific kind of validation experiment. Validation experiments are broader: they can be a customer interview, a landing page, a concierge test, or any of six build methods. Every MVP is a validation experiment; not every validation experiment produces an MVP.

What counts as evidence in startup validation?

Observable customer behavior, weighted by the cost the customer paid to produce it. Stronger signals are commitment (deposits, pre-orders, paid pilots), use (return visits, completed workflows, retention at the natural frequency), and referral (unprompted word-of-mouth, customer-to-customer defense). Weaker signals are page views without intent, compliments, survey interest, and social likes from outside the customer segment.

What are vanity metrics in startup validation?

Numbers that look impressive but cannot change a decision. Eric Ries distinguishes them from actionable metrics, which tie specific behaviors to observed outcomes. Examples include total signups without qualification, page views without intent, beta-tester compliments, and social likes from outside the segment. The signal is real (something happened); the implication is weak (nothing was learned that changes a decision).

When should a founder change direction?

When the loop produces evidence that the riskiest assumption is unlikely to hold. The decision branches are continue, change the assumption, change the customer, change the solution, change the channel, and stop. Pivoting is a structured course correction — the founder names the new hypothesis, designs the next experiment, and runs the loop again with the new assumption. Panic is not pivoting.

Can validation guarantee startup success?

No. Lean Startup validation reduces uncertainty; it does not predict outcomes. The honest framing is empirical: the loop narrows the range of things the founder could still be wrong about, which is the only thing validation can honestly do.

Summary

Lean Startup validation is the discipline of testing startup assumptions through iterative Build-Measure-Learn experiments before committing to a full build. Eric Ries's The Lean Startup (2011) defined the loop as a way to turn assumptions into validated knowledge. The loop is small, but choosing the next experiment is harder than running it.

The five-question framework picks the next experiment: what is the assumption, how uncertain are you, what behavior would change your mind, what is the cheapest credible experiment, and what result would force a new decision. The four decision thresholds — useful evidence, inconclusive, re-run trigger, reconsider trigger — are written before the experiment runs. The six named decisions are continue, change the assumption, change the customer, change the solution, change the channel, and stop.

The smallest "build" can be a landing page, an interview, a concierge MVP, a Wizard-of-Oz service, a prototype, a paid pilot, or a small feature-scope MVP. The right method matches the assumption.

The discipline is not a verdict on success. It is a method for narrowing the range of things the founder could still be wrong about. The honest framing is empirical: the loop produces evidence, not certainty.

What to do next

The smallest useful action this week is to write down the riskiest assumption in the format "We believe [customer] will [behavior] when [condition]." The next is to answer the five questions in the framework above and turn the assumption into a one-page experiment design. The third is to run the experiment with a specific segment for a specific duration, measure observable behavior, and write the learning in one sentence.

If you want a structured second opinion on which assumption the loop should test first, Startup MRI's validation analysis helps — it returns the critical assumption your idea most depends on so the experiment you run is the one with the highest signal. For worked examples of what a real validation report looks like, see the SaaS validation report example, the AI startup validation report example, or the mobile app validation report example.

Run the loop. Write the learning. Decide what to test next.

Test your own idea

Describe your idea, answer five short questions, and get a structured 8-dimension report — free, no signup.

Continue learning

Where to go from here

These pieces are grouped by topic, not publication date — pick the one that matches the question you are working on right now.

More in ValidationSee all topics →

See every article on startup validation in one place.

Open the Startup Validation hub →