Skip to content
YibudYibud

How we validate

Startup MRI Methodology

Startup MRI treats validation as a learning process, not a prediction. The goal is to make the riskiest assumptions visible before you commit to building.

Last updated · September 1, 2026

Quick answer

What is Startup MRI methodology?

It is an evidence-first framework for examining whether a problem is real, a customer is reachable, a market can support a business, and a proposed way to earn revenue is testable. The report organizes signals and tradeoffs; it does not promise that an idea will succeed.

Key principles

Five questions before you build

Evidence before assumptions

Write down what must be true, then look for observable behavior rather than relying on enthusiasm or intuition.

Customer validation

Talk with people who experience the problem and ask about their recent workarounds, costs, and decisions.

Market validation

Study the alternatives customers already use and the constraints that shape the market you can actually reach.

Distribution validation

Treat access to customers as a hypothesis. A useful product still needs a credible path to its first users.

Monetization validation

Separate compliments from commitment by testing a concrete price, offer, or paid pilot when appropriate.

Validation framework

From idea to decision support

  1. 1

    Idea

    Describe the customer, problem, and proposed change in plain language.

  2. 2

    Assumptions

    Identify the beliefs about demand, competition, distribution, revenue, and execution that could make or break the idea.

  3. 3

    Validation signals

    Choose observable tests: problem interviews, existing behavior, landing-page responses, waitlists, pilots, or payment conversations.

  4. 4

    Risk assessment

    Compare the strength of signals with the consequences if an assumption is wrong. A weak signal is not a verdict; it is a reason to learn more. The same call surfaces a single highest-risk dimension used to derive a 7-Day Validation Action Plan.

  5. 5

    Decision support

    Use the findings to narrow an MVP, choose the next experiment, or stop and revise the idea before investing further. The report ships with a 7-Day Validation Action Plan — one day at a time, with the evidence to collect and a Day-7 Continue / Refine / Re-test / Stop signal.

Inside the score

Seven dimensions, weighted on purpose

Every number in a report comes from a deterministic rule engine, not from the language model. The engine reads your idea text and your five structured answers, fires whichever of its rules match, and scores seven dimensions from 0 to 100. The language model then explains those scores. It never sets them, which is why the same inputs always produce the same report. What follows is the part most tools keep to themselves: which dimensions exist, what each one actually asks, and how much each one counts.

  • Validation

    25%

    How cheaply and concretely can this be tested? A specific, named audience and an idea that already contains numbers or pilot results score higher. An idea needing regulated or physical proof — hardware, clinical, licensed — scores lower, because the first honest experiment is expensive, not because the idea is weak.

  • Distribution

    20%

    Is there a credible path to the first hundred customers? Community-led and niche-forum motions score highest for a solo founder. Paid acquisition and cold outreach score negative, because both tend to stall before a product has proof. Arriving with no chosen channel is the heaviest penalty in this dimension.

  • Market demand

    18%

    Is the audience specific enough to find, and is the problem described concretely enough to check? Vague audiences — “everyone”, “users” — and one-line ideas lose points here. Not because the idea is bad, but because nothing in the input can be tested yet.

  • Monetization

    12%

    Is there a revenue model you can test at a real price? Recurring subscription scores highest. Marketplaces are treated as neutral, because of the chicken-and-egg problem rather than any doubt about the model. Arriving without a monetization answer is the single heaviest input penalty anywhere in the engine.

  • Competition

    10%

    Is the category already crowded with near-identical products, and is the audience narrow enough to defend? Ideas matching the patterns of a saturated category lose points here.

  • Founder skills

    10%

    Does your stated background match what shipping and selling this would require? An experienced developer gains the most. A non-technical founder loses a small amount — hiring or learning is a real cost, but not a disqualification.

  • Build ease

    5%

    How much work stands between the idea and something a customer can react to? Technical background is the only input that moves this dimension.

Why validation carries the most weight

The seven weights sum to 1.00, and validation holds the largest share at 25 per cent. That is an opinion rather than a neutral average, and it is worth saying plainly: the engine assumes the most likely reason a project dies is that nobody checked. Build ease carries the least, at 5 per cent. It used to carry twice that. It was cut because writing software has become the cheap part of starting a company, while finding out whether anyone wants it has not. If you disagree with that ordering, the report still works. Every dimension is scored and shown separately, so you can set the weighting aside and read the breakdown on your own terms.

Two numbers that mean different things

Reports carry an overall startup score and an opportunity score. They are computed differently on purpose. The overall score is the weighted sum above, so it carries the engine's priorities. The opportunity score is the plain average of the same seven dimensions, each counting equally. When the two diverge, the gap is itself the signal. Scoring higher on opportunity than overall usually means you are strong where the engine assigns little weight and weak where it assigns a lot — most often comfortable on build and founder skills, thin on validation and distribution.

How scores become a verdict

The verdict is a threshold rule, and each half does separate work. Go needs an overall score of at least 75 and no single dimension below 50. Caution needs an overall score of at least 55 and no single dimension below 30. Everything else is No-go. The second clause in each rule is the one that matters. An idea can average comfortably and still be rejected because one dimension is critically weak — the classic shape being a genuinely good product with no plausible way to reach anybody.

What a score is not

A score measures how completely and specifically your inputs address a dimension. It does not measure your idea, and it is not a probability of success. The engine has never met your customers. The distinction is practical. A low distribution score does not mean distribution is impossible; it means the channel you named does not obviously reach the audience you named, so that is now the cheapest thing to go test. A high score does not mean build it; it means the engine found no stated reason not to, given what you told it. When a score looks wrong it is almost always traceable to one input. Change that input and the score changes deterministically. That traceability is the entire reason the scores are not left to a language model.

Worked example

From inputs to a verdict on a concrete idea

What follows is an illustrative example. The numbers come from the same deterministic rule engine the live product runs; the names, the customers, and the waitlist size are invented for explanation. The fired rules and resulting scores are real.

Illustrative example. No real company is referenced. Numbers shown are plausible operating points, not verified benchmarks.

The scenario

A solo founder is considering an SMS tool for Shopify stores

The founder has already run 25 problem interviews, signed 8 letters of intent, and collected 140 emails on a waitlist. The plan is to charge a recurring monthly fee and reach merchants through the Shopify merchant community (Slack groups, two subreddits, one specialist newsletter). The next question is not whether the idea is interesting — the waitlist says that — but what the score is, where the score comes from, and what to test before building.

Step 1 — the input

What the founder typed into the Analyzer

Five structured fields plus the idea text. The rule engine reads exactly these. Nothing else.

Startup idea
A B2B SaaS that helps Shopify stores doing $50K monthly revenue recover abandoned carts through personalised SMS. We interviewed 25 store owners, signed 8 LOIs, and have a waitlist of 140.
Target audience
Shopify stores doing $50K monthly revenue with 5K-20K monthly visitors
Monetization model
Subscription
Acquisition channel
Community (Shopify merchant Slack groups, two subreddits, one newsletter)
Technical background
Experienced developer

Step 2 — what the engine sees

Which rules fire on these inputs

Each rule is a pattern that matches against the inputs and adds or subtracts a fixed adjustment on the dimensions it affects. Seven rules fire for this idea. None subtract.

RuleWhy it firesAdjustment
mon.subscriptionSubscription is a proven, recurring model.+10 to monetization, competition
acq.communityCommunity-led growth is defensible and high-trust.+7 to distribution
bg.experiencedAn experienced developer can ship the MVP solo.+5 to founder fit, build
idea.detailedLonger, more specific ideas tend to be better thought through.+3 to market
val.cost_lowValidation can be done cheaply via a landing page or waitlist.+3 to validation
audience.icp_clearTarget audience is specific (named role, industry, revenue, or job-to-be-done).+4 to market, validation
evidence.denseIdea includes specific numbers, data, or pilot evidence.+3 to market, validation

Step 3 — per-dimension score

Base + adjustments, clamped to 0–100

Every dimension starts at a baseline. Each rule that affects a dimension adds or subtracts. The result is clamped to a 0–100 integer.

DimensionBaseAdjustmentsScore
Market45+3, +4, +355
Competition50+1060
Distribution40+747
Monetization45+1055
Build55+560
Founder fit50+555
Validation50+3, +4, +360

Step 4 — two summary numbers

Overall vs opportunity

The overall score is the weighted sum across all seven dimensions; the opportunity score is the plain average. When they diverge, the gap is itself the signal. When they sit close together, no dimension is masking another.

Overall score (weighted)

55

0.25·validation + 0.20·distribution + 0.18·market + 0.12·monetization + 0.10·competition + 0.10·founder_fit + 0.05·build = 55.4 → 55

Opportunity score (plain average)

56

(55 + 60 + 47 + 55 + 60 + 55 + 60) ÷ 7 = 56.0

Gap = 1

The two numbers sit almost on top of each other. There is no dimension where the engine assigns little weight but the founder is strong, and no dimension where the engine assigns high weight but the founder is weak. The input is balanced.

Step 5 — verdict

How the threshold rule reads these scores

The rule

Go requires overall ≥ 75 AND no dimension below 50. Caution requires overall ≥ 55 AND no dimension below 30. Everything else is No-go.

Applied to this idea

  • Go: overall 55 is below 75 — not met.
  • Caution: overall 55 meets the minimum, and the lowest dimension is 47 (≥ 30) — met.

Verdict: Caution.

What this means

The structure of the input — sharp ICP, dense evidence, named channel, recurring monetization — already does a lot of work. Validation lands at 60 because the input contains a specific waitlist size and interview count, not because the product is good. Distribution sits at 47 because community is a defensible first channel but does not compound on its own. There is no dimension critically weak (all ≥ 30) and no dimension strong enough to clear the Go bar (all < 75). The decision is not “build now.” It is “test the willingness-to-pay and the renewal signal before expanding the build.”

Step 6 — next experiment

The cheapest experiment that targets the remaining risk

The waitlist is already evidence of stated intent. The remaining unknown is whether those merchants will pay the recurring price and stay. The cheapest experiment that tests that — without expanding the build — is a 30-day paid pilot with five waitlist members at the planned price (for example $99/month). The pilot’s job is not to validate interest. The waitlist already did that. The pilot’s job is to test renewal. Without a renewal signal, the waitlist is still evidence of interest, not of payment.

What this does not prove

A 55 caution is not a verdict on the idea. It is a structured read on what the inputs say. The engine has not met the customers, has not seen the LOIs, and has not priced the SMS carrier. The same idea with a different acquisition channel would score differently. Switch the channel from community to “not sure” and distribution drops from 47 to 34; overall drops to 53. The change is traceable to a single input. That traceability is the entire reason the scores are not left to a language model.

See the same shape in a full report

The reasoning above is the engine working in the open. The example reports show what the same engine produces once the AI explanation layer is added on top of those scores.

Why this matters

Validation buys learning, not certainty

Founders validate because building is expensive and early confidence is often based on untested assumptions. A structured process helps turn a broad idea into a short list of questions that can be answered through customer behavior. Startup MRI makes that list easier to inspect, while real conversations and experiments remain the source of truth.

What this methodology cannot do

AI cannot predict startup success, replace customer research, or know facts that were never provided. Scores are decision-support signals produced from structured inputs, not probabilities of success. Treat the report as a starting point for experiments and update your view when new evidence disagrees.

Evidence re-evaluation

What evidence can change in a re-evaluation

After you run the Recommended Experiment from the report, you can submit what happened on the report page itself. Startup MRI runs the same deterministic engine against your new evidence and surfaces a side-by-side comparison: original assessment, evidence you submitted, updated assessment, what changed, and the next recommended action. The model is grounded in published customer-development and lean-startup literature (see Sources). It does not invent offsets; every change is conservative and tied to the kind of evidence you collected.

Evidence strength is classified, not counted

Evidence is classified into three tiers by kind, not by sample size alone. Weak evidence — polite interest, hypothetical answers, founder interpretation without behavioural confirmation. Does not change the assessment. Moderate evidence — multiple independent descriptions of the same recent problem, observed workaround behaviour, quantified stated intent. Updates the explanation. Stronger evidence — prepayment, signed LOI, quantified conversion at a non-trivial rate, completed prototype sessions. Can reduce or sharpen uncertainty on the dimension the experiment targeted. More interviews alone is not stronger evidence; behavioural evidence at small N is.

Why more interviews do not automatically mean a higher score

Three reasons. First, polite interest is not commitment. The Mom Test (Fitzpatrick) classifies compliments as the fool's gold of customer learning: shiny, distracting, and worthless. We treat them as Weak evidence and they do not move scores. Second, quantity alone is insufficient. Strategyzer's Testing Business Ideas framework distinguishes behavioural evidence (observed action) from stated evidence (verbal intent) and from opinion (founder interpretation). A hundred polite yeses are still stated evidence; three paying customers are behavioural. Third, the original assessment is not invalidated by positive feedback on one experiment. The model requires the evidence to clear three gates — experiment-to-dimension match, at-least-Moderate tier, and consistency — before any uncertainty status flips. A single experiment that passes all three changes a small amount.

What evidence re-evaluation cannot do

It cannot predict startup success. It cannot lift a dimension the experiment did not target. It cannot make Weak evidence into Stronger evidence by collecting more of it. It cannot override the original engine rules — those remain the authoritative voice on what the original input implies. It is also conservative by design. The largest positive offset on any single dimension is +8 (less than the engine's typical single-rule contribution). The largest negative offset is -12, slightly larger than positive, because negative evidence is harder to fake. A single re-evaluation cannot flip the overall verdict unless at least two dimensions changed OR a Stronger + Moderate pair on the same dimension agrees. Re-evaluation is a learning loop, not a green light.

Common questions

Startup validation methodology FAQ

How does startup validation work?

Start with assumptions, identify the riskiest one, and run a small test that can produce a believable no as well as a yes.

What counts as evidence?

Recent behavior, specific past experiences, repeated workarounds, trial commitments, and payment or pilot decisions are stronger than hypothetical praise.

Why speak with customers before building?

Customer conversations can reveal the language, urgency, existing alternatives, and context that a founder cannot infer from an idea alone.

What is market validation?

It is testing whether a reachable group has a meaningful problem and whether the surrounding alternatives leave room for a useful offer.

Why is distribution part of validation?

A product idea is incomplete until you can describe how the intended customer will discover and adopt it.

How should I test monetization?

Ask for a concrete commitment such as a paid pilot, deposit, or purchase conversation instead of treating stated interest as willingness to pay.

Do the scores predict success?

No. They summarize structured inputs and tradeoffs to support a decision about what to test next.

Can AI predict startup success?

No. Startup outcomes depend on changing markets, execution, customer behavior, and factors a report cannot observe.

Turn your assumptions into a next step

Describe your idea and get a structured starting point for validation.

Analyze my idea →