How we validate
Startup MRI Methodology
Startup MRI treats validation as a learning process, not a prediction. The goal is to make the riskiest assumptions visible before you commit to building.
Last updated · September 1, 2026
Quick answer
What is Startup MRI methodology?
It is an evidence-first framework for examining whether a problem is real, a customer is reachable, a market can support a business, and a proposed way to earn revenue is testable. The report organizes signals and tradeoffs; it does not promise that an idea will succeed.
Key principles
Five questions before you build
Evidence before assumptions
Write down what must be true, then look for observable behavior rather than relying on enthusiasm or intuition.
Customer validation
Talk with people who experience the problem and ask about their recent workarounds, costs, and decisions.
Market validation
Study the alternatives customers already use and the constraints that shape the market you can actually reach.
Distribution validation
Treat access to customers as a hypothesis. A useful product still needs a credible path to its first users.
Monetization validation
Separate compliments from commitment by testing a concrete price, offer, or paid pilot when appropriate.
Validation framework
From idea to decision support
- 1
Idea
Describe the customer, problem, and proposed change in plain language.
- 2
Assumptions
Identify the beliefs about demand, competition, distribution, revenue, and execution that could make or break the idea.
- 3
Validation signals
Choose observable tests: problem interviews, existing behavior, landing-page responses, waitlists, pilots, or payment conversations.
- 4
Risk assessment
Compare the strength of signals with the consequences if an assumption is wrong. A weak signal is not a verdict; it is a reason to learn more. The same call surfaces a single highest-risk dimension used to derive a 7-Day Validation Action Plan.
- 5
Decision support
Use the findings to narrow an MVP, choose the next experiment, or stop and revise the idea before investing further. The report ships with a 7-Day Validation Action Plan — one day at a time, with the evidence to collect and a Day-7 Continue / Refine / Re-test / Stop signal.
Inside the score
Seven dimensions, weighted on purpose
Every number in a report comes from a deterministic rule engine, not from the language model. The engine reads your idea text and your five structured answers, fires whichever of its rules match, and scores seven dimensions from 0 to 100. The language model then explains those scores. It never sets them, which is why the same inputs always produce the same report. What follows is the part most tools keep to themselves: which dimensions exist, what each one actually asks, and how much each one counts.
Validation
25%How cheaply and concretely can this be tested? A specific, named audience and an idea that already contains numbers or pilot results score higher. An idea needing regulated or physical proof — hardware, clinical, licensed — scores lower, because the first honest experiment is expensive, not because the idea is weak.
Distribution
20%Is there a credible path to the first hundred customers? Community-led and niche-forum motions score highest for a solo founder. Paid acquisition and cold outreach score negative, because both tend to stall before a product has proof. Arriving with no chosen channel is the heaviest penalty in this dimension.
Market demand
18%Is the audience specific enough to find, and is the problem described concretely enough to check? Vague audiences — “everyone”, “users” — and one-line ideas lose points here. Not because the idea is bad, but because nothing in the input can be tested yet.
Monetization
12%Is there a revenue model you can test at a real price? Recurring subscription scores highest. Marketplaces are treated as neutral, because of the chicken-and-egg problem rather than any doubt about the model. Arriving without a monetization answer is the single heaviest input penalty anywhere in the engine.
Competition
10%Is the category already crowded with near-identical products, and is the audience narrow enough to defend? Ideas matching the patterns of a saturated category lose points here.
Founder skills
10%Does your stated background match what shipping and selling this would require? An experienced developer gains the most. A non-technical founder loses a small amount — hiring or learning is a real cost, but not a disqualification.
Build ease
5%How much work stands between the idea and something a customer can react to? Technical background is the only input that moves this dimension.
Why validation carries the most weight
The seven weights sum to 1.00, and validation holds the largest share at 25 per cent. That is an opinion rather than a neutral average, and it is worth saying plainly: the engine assumes the most likely reason a project dies is that nobody checked. Build ease carries the least, at 5 per cent. It used to carry twice that. It was cut because writing software has become the cheap part of starting a company, while finding out whether anyone wants it has not. If you disagree with that ordering, the report still works. Every dimension is scored and shown separately, so you can set the weighting aside and read the breakdown on your own terms.
Two numbers that mean different things
Reports carry an overall startup score and an opportunity score. They are computed differently on purpose. The overall score is the weighted sum above, so it carries the engine's priorities. The opportunity score is the plain average of the same seven dimensions, each counting equally. When the two diverge, the gap is itself the signal. Scoring higher on opportunity than overall usually means you are strong where the engine assigns little weight and weak where it assigns a lot — most often comfortable on build and founder skills, thin on validation and distribution.
How scores become a verdict
The verdict is a threshold rule, and each half does separate work. Go needs an overall score of at least 75 and no single dimension below 50. Caution needs an overall score of at least 55 and no single dimension below 30. Everything else is No-go. The second clause in each rule is the one that matters. An idea can average comfortably and still be rejected because one dimension is critically weak — the classic shape being a genuinely good product with no plausible way to reach anybody.
What a score is not
A score measures how completely and specifically your inputs address a dimension. It does not measure your idea, and it is not a probability of success. The engine has never met your customers. The distinction is practical. A low distribution score does not mean distribution is impossible; it means the channel you named does not obviously reach the audience you named, so that is now the cheapest thing to go test. A high score does not mean build it; it means the engine found no stated reason not to, given what you told it. When a score looks wrong it is almost always traceable to one input. Change that input and the score changes deterministically. That traceability is the entire reason the scores are not left to a language model.
Worked example
From inputs to a verdict on a concrete idea
What follows is an illustrative example. The numbers come from the same deterministic rule engine the live product runs; the names, the customers, and the waitlist size are invented for explanation. The fired rules and resulting scores are real.
Illustrative example. No real company is referenced. Numbers shown are plausible operating points, not verified benchmarks.
The scenario
A solo founder is considering an SMS tool for Shopify stores
The founder has already run 25 problem interviews, signed 8 letters of intent, and collected 140 emails on a waitlist. The plan is to charge a recurring monthly fee and reach merchants through the Shopify merchant community (Slack groups, two subreddits, one specialist newsletter). The next question is not whether the idea is interesting — the waitlist says that — but what the score is, where the score comes from, and what to test before building.
Step 1 — the input
What the founder typed into the Analyzer
Five structured fields plus the idea text. The rule engine reads exactly these. Nothing else.
- Startup idea
- A B2B SaaS that helps Shopify stores doing $50K monthly revenue recover abandoned carts through personalised SMS. We interviewed 25 store owners, signed 8 LOIs, and have a waitlist of 140.
- Target audience
- Shopify stores doing $50K monthly revenue with 5K-20K monthly visitors
- Monetization model
- Subscription
- Acquisition channel
- Community (Shopify merchant Slack groups, two subreddits, one newsletter)
- Technical background
- Experienced developer
Step 2 — what the engine sees
Which rules fire on these inputs
Each rule is a pattern that matches against the inputs and adds or subtracts a fixed adjustment on the dimensions it affects. Seven rules fire for this idea. None subtract.
| Rule | Why it fires | Adjustment |
|---|---|---|
| mon.subscription | Subscription is a proven, recurring model. | +10 to monetization, competition |
| acq.community | Community-led growth is defensible and high-trust. | +7 to distribution |
| bg.experienced | An experienced developer can ship the MVP solo. | +5 to founder fit, build |
| idea.detailed | Longer, more specific ideas tend to be better thought through. | +3 to market |
| val.cost_low | Validation can be done cheaply via a landing page or waitlist. | +3 to validation |
| audience.icp_clear | Target audience is specific (named role, industry, revenue, or job-to-be-done). | +4 to market, validation |
| evidence.dense | Idea includes specific numbers, data, or pilot evidence. | +3 to market, validation |
Step 3 — per-dimension score
Base + adjustments, clamped to 0–100
Every dimension starts at a baseline. Each rule that affects a dimension adds or subtracts. The result is clamped to a 0–100 integer.
| Dimension | Base | Adjustments | Score |
|---|---|---|---|
| Market | 45 | +3, +4, +3 | 55 |
| Competition | 50 | +10 | 60 |
| Distribution | 40 | +7 | 47 |
| Monetization | 45 | +10 | 55 |
| Build | 55 | +5 | 60 |
| Founder fit | 50 | +5 | 55 |
| Validation | 50 | +3, +4, +3 | 60 |
Step 4 — two summary numbers
Overall vs opportunity
The overall score is the weighted sum across all seven dimensions; the opportunity score is the plain average. When they diverge, the gap is itself the signal. When they sit close together, no dimension is masking another.
Overall score (weighted)
55
0.25·validation + 0.20·distribution + 0.18·market + 0.12·monetization + 0.10·competition + 0.10·founder_fit + 0.05·build = 55.4 → 55
Opportunity score (plain average)
56
(55 + 60 + 47 + 55 + 60 + 55 + 60) ÷ 7 = 56.0
Gap = 1
The two numbers sit almost on top of each other. There is no dimension where the engine assigns little weight but the founder is strong, and no dimension where the engine assigns high weight but the founder is weak. The input is balanced.
Step 5 — verdict
How the threshold rule reads these scores
The rule
Go requires overall ≥ 75 AND no dimension below 50. Caution requires overall ≥ 55 AND no dimension below 30. Everything else is No-go.
Applied to this idea
- Go: overall 55 is below 75 — not met.
- Caution: overall 55 meets the minimum, and the lowest dimension is 47 (≥ 30) — met.
Verdict: Caution.
What this means
The structure of the input — sharp ICP, dense evidence, named channel, recurring monetization — already does a lot of work. Validation lands at 60 because the input contains a specific waitlist size and interview count, not because the product is good. Distribution sits at 47 because community is a defensible first channel but does not compound on its own. There is no dimension critically weak (all ≥ 30) and no dimension strong enough to clear the Go bar (all < 75). The decision is not “build now.” It is “test the willingness-to-pay and the renewal signal before expanding the build.”
Step 6 — next experiment
The cheapest experiment that targets the remaining risk
The waitlist is already evidence of stated intent. The remaining unknown is whether those merchants will pay the recurring price and stay. The cheapest experiment that tests that — without expanding the build — is a 30-day paid pilot with five waitlist members at the planned price (for example $99/month). The pilot’s job is not to validate interest. The waitlist already did that. The pilot’s job is to test renewal. Without a renewal signal, the waitlist is still evidence of interest, not of payment.
What this does not prove
A 55 caution is not a verdict on the idea. It is a structured read on what the inputs say. The engine has not met the customers, has not seen the LOIs, and has not priced the SMS carrier. The same idea with a different acquisition channel would score differently. Switch the channel from community to “not sure” and distribution drops from 47 to 34; overall drops to 53. The change is traceable to a single input. That traceability is the entire reason the scores are not left to a language model.
See the same shape in a full report
The reasoning above is the engine working in the open. The example reports show what the same engine produces once the AI explanation layer is added on top of those scores.
Why this matters
Validation buys learning, not certainty
Founders validate because building is expensive and early confidence is often based on untested assumptions. A structured process helps turn a broad idea into a short list of questions that can be answered through customer behavior. Startup MRI makes that list easier to inspect, while real conversations and experiments remain the source of truth.
What this methodology cannot do
AI cannot predict startup success, replace customer research, or know facts that were never provided. Scores are decision-support signals produced from structured inputs, not probabilities of success. Treat the report as a starting point for experiments and update your view when new evidence disagrees.
Evidence re-evaluation
What evidence can change in a re-evaluation
After you run the Recommended Experiment from the report, you can submit what happened on the report page itself. Startup MRI runs the same deterministic engine against your new evidence and surfaces a side-by-side comparison: original assessment, evidence you submitted, updated assessment, what changed, and the next recommended action. The model is grounded in published customer-development and lean-startup literature (see Sources). It does not invent offsets; every change is conservative and tied to the kind of evidence you collected.
Evidence strength is classified, not counted
Evidence is classified into three tiers by kind, not by sample size alone. Weak evidence — polite interest, hypothetical answers, founder interpretation without behavioural confirmation. Does not change the assessment. Moderate evidence — multiple independent descriptions of the same recent problem, observed workaround behaviour, quantified stated intent. Updates the explanation. Stronger evidence — prepayment, signed LOI, quantified conversion at a non-trivial rate, completed prototype sessions. Can reduce or sharpen uncertainty on the dimension the experiment targeted. More interviews alone is not stronger evidence; behavioural evidence at small N is.
Why more interviews do not automatically mean a higher score
Three reasons. First, polite interest is not commitment. The Mom Test (Fitzpatrick) classifies compliments as the fool's gold of customer learning: shiny, distracting, and worthless. We treat them as Weak evidence and they do not move scores. Second, quantity alone is insufficient. Strategyzer's Testing Business Ideas framework distinguishes behavioural evidence (observed action) from stated evidence (verbal intent) and from opinion (founder interpretation). A hundred polite yeses are still stated evidence; three paying customers are behavioural. Third, the original assessment is not invalidated by positive feedback on one experiment. The model requires the evidence to clear three gates — experiment-to-dimension match, at-least-Moderate tier, and consistency — before any uncertainty status flips. A single experiment that passes all three changes a small amount.
What evidence re-evaluation cannot do
It cannot predict startup success. It cannot lift a dimension the experiment did not target. It cannot make Weak evidence into Stronger evidence by collecting more of it. It cannot override the original engine rules — those remain the authoritative voice on what the original input implies. It is also conservative by design. The largest positive offset on any single dimension is +8 (less than the engine's typical single-rule contribution). The largest negative offset is -12, slightly larger than positive, because negative evidence is harder to fake. A single re-evaluation cannot flip the overall verdict unless at least two dimensions changed OR a Stronger + Moderate pair on the same dimension agrees. Re-evaluation is a learning loop, not a green light.
Common questions
Startup validation methodology FAQ
How does startup validation work?
Start with assumptions, identify the riskiest one, and run a small test that can produce a believable no as well as a yes.
What counts as evidence?
Recent behavior, specific past experiences, repeated workarounds, trial commitments, and payment or pilot decisions are stronger than hypothetical praise.
Why speak with customers before building?
Customer conversations can reveal the language, urgency, existing alternatives, and context that a founder cannot infer from an idea alone.
What is market validation?
It is testing whether a reachable group has a meaningful problem and whether the surrounding alternatives leave room for a useful offer.
Why is distribution part of validation?
A product idea is incomplete until you can describe how the intended customer will discover and adopt it.
How should I test monetization?
Ask for a concrete commitment such as a paid pilot, deposit, or purchase conversation instead of treating stated interest as willingness to pay.
Do the scores predict success?
No. They summarize structured inputs and tradeoffs to support a decision about what to test next.
Can AI predict startup success?
No. Startup outcomes depend on changing markets, execution, customer behavior, and factors a report cannot observe.
Turn your assumptions into a next step
Describe your idea and get a structured starting point for validation.
Analyze my idea →