验证方法
Startup MRI 验证方法
Startup MRI treats validation as a learning process, not a prediction. The goal is to make the riskiest assumptions visible before you commit to building.
Last updated · 2026年9月1日
快速回答
Startup MRI 的方法是什么?
It is an evidence-first framework for examining whether a problem is real, a customer is reachable, a market can support a business, and a proposed way to earn revenue is testable. The report organizes signals and tradeoffs; it does not promise that an idea will succeed.
核心原则
构建之前先回答五个问题
Evidence before assumptions
Write down what must be true, then look for observable behavior rather than relying on enthusiasm or intuition.
Customer validation
Talk with people who experience the problem and ask about their recent workarounds, costs, and decisions.
Market validation
Study the alternatives customers already use and the constraints that shape the market you can actually reach.
Distribution validation
Treat access to customers as a hypothesis. A useful product still needs a credible path to its first users.
Monetization validation
Separate compliments from commitment by testing a concrete price, offer, or paid pilot when appropriate.
验证框架
从想法到决策支持
- 1
Idea
Describe the customer, problem, and proposed change in plain language.
- 2
Assumptions
Identify the beliefs about demand, competition, distribution, revenue, and execution that could make or break the idea.
- 3
Validation signals
Choose observable tests: problem interviews, existing behavior, landing-page responses, waitlists, pilots, or payment conversations.
- 4
Risk assessment
Compare the strength of signals with the consequences if an assumption is wrong. A weak signal is not a verdict; it is a reason to learn more. The same call surfaces a single highest-risk dimension used to derive a 7-Day Validation Action Plan.
- 5
Decision support
Use the findings to narrow an MVP, choose the next experiment, or stop and revise the idea before investing further. The report ships with a 7-Day Validation Action Plan — one day at a time, with the evidence to collect and a Day-7 Continue / Refine / Re-test / Stop signal.
分数的内部结构
七个维度,权重是刻意设定的
报告里的每一个数字都来自确定性规则引擎,而不是语言模型。引擎读取你的想法描述和五项结构化回答,触发匹配的规则,对七个维度各给出 0 到 100 的分数。语言模型只负责解释这些分数,从不决定它们——这也是相同输入永远得到相同报告的原因。 下面是多数同类工具不会公开的部分:有哪些维度、每个维度实际在问什么、以及各自占多少权重。
验证可行性
25%这件事能多便宜、多具体地被验证?受众明确具体、想法里已经带有数字或试点结果的,得分更高。需要监管或实物验证的想法(硬件、临床、牌照)得分更低——原因是第一个诚实的实验成本很高,而不是想法本身不好。
分发渠道
20%有没有触达前一百个客户的可信路径?对独立开发者而言,社区驱动和垂直论坛得分最高。付费投放和陌生外拓为负分,因为在产品还没有证据之前,这两条路通常走不动。完全没有选定渠道,是本维度扣分最重的情形。
市场需求
18%受众是否具体到找得到?问题是否描述得足够具体、可以核查?模糊受众(「所有人」「用户」)和一句话想法会在这里失分——不是因为想法差,而是因为输入里还没有任何东西可供验证。
变现模式
12%有没有一个能用真实价格去测试的收入模式?订阅制得分最高。平台型被视为中性,原因是先有鸡还是先有蛋的问题,而不是对该模式本身有疑虑。没有给出变现答案,是整个引擎中最重的单项输入扣分。
竞争格局
10%这个品类是否已挤满高度相似的产品?受众是否窄到足以守住?命中饱和品类特征的想法会在这里失分。
创始人匹配度
10%你申明的技术背景,是否匹配把这件事做出来并卖出去所需的能力?有经验的开发者加分最多。非技术创始人小幅扣分——招人或自学是真实成本,但不构成否决。
构建难度
5%从想法到客户能够真实反馈的东西之间,隔着多少工作量?技术背景是唯一影响这个维度的输入。
为什么验证可行性权重最高
七项权重之和为 1.00,其中验证可行性占比最大,为 25%。这是一种观点,而不是中立的平均分配,也值得直说:引擎假定一个项目最可能的死因,是没有人去核实过。 构建难度权重最低,只有 5%,此前是这个数的两倍。下调的理由是:写软件已经变成创业里便宜的那部分,而搞清楚有没有人想要它,并没有变便宜。 如果你不同意这个排序,报告依然可用。每个维度都单独计分、单独展示,你完全可以把权重放在一边,按自己的标准读这份拆解。
两个数字,含义不同
报告同时给出总体创业分和机会分,两者的算法是刻意不同的。 总体分是上面那组权重的加权和,因此它承载了引擎的优先级。机会分是同样七个维度的简单平均,每项等权。 当两者出现差距时,这个差距本身就是信号。机会分高于总体分,通常意味着你在引擎赋予低权重的地方很强、在赋予高权重的地方很弱——最常见的组合是构建和创始人匹配度都不错,但验证和分发很薄。
分数如何变成结论
结论由阈值规则给出,每条规则的两半各有分工。 Go 需要总体分不低于 75,且没有任何单项低于 50。 Caution 需要总体分不低于 55,且没有任何单项低于 30。 其余一切都是 No-go。 真正起作用的是每条规则的后半句。一个想法完全可以平均分不错却仍被否决,因为某一个维度已经低到危险区——最典型的形态,就是产品确实不错,但没有任何可行办法触达到人。
分数不是什么
分数衡量的是你的输入对某个维度回答得多完整、多具体。它不衡量你的想法,也不是成功概率。引擎从未见过你的客户。 这个区分很实用。分发分低不代表分发做不成,而是说你填的渠道并不明显能触达你填的受众,所以这件事现在是最便宜的待验证项。分数高也不代表就该开始做,只代表基于你提供的信息,引擎没找到明确的反对理由。 当某个分数看起来不对,几乎总能追溯到某一项输入。改掉那项输入,分数会确定性地改变。正是这种可追溯性,才使得这些分数不交给语言模型来决定。
实例演练
从输入到对一个具体想法的结论
以下是一个说明性示例。数字来自产品实际运行的同一套确定性规则引擎;其中的人名、客户名和候补名单规模都是为了讲解而虚构的。被触发的规则和得出的分数是真实的。
说明性示例。并未引用任何真实公司。所列数字是合理的运行区间,而非经验证的基准。
场景
一位独立开发者正在评估一款面向 Shopify 卖家的短信挽回工具
这位开发者已经完成了 25 场问题访谈,签署了 8 份意向书,并积累了 140 个候补名单邮箱。计划按月订阅收费,通过 Shopify 商家社群(Slack 群、两个 Reddit 社区、一份垂直邮件简报)触达客户。下一个问题不是这个想法是否有趣——候补名单已经回答了——而是分数是多少、分数从哪来、在动手做之前还应该测什么。
第一步 — 输入
开发者在分析器里填写的内容
五项结构化字段加上想法描述。规则引擎只读取这些信息,不读其他。
- 创业想法
- 一款 B2B SaaS,帮助月营收 5 万美元的 Shopify 独立站通过个性化短信挽回弃单。我们已经访谈了 25 位店主,签署了 8 份意向书,候补名单有 140 人。
- 目标受众
- 月营收 5 万美元、月访问量 5000–20000 的 Shopify 独立站
- 变现模式
- 订阅制
- 获客渠道
- 社群(Shopify 商家 Slack 群、两个 Reddit 社区、一份邮件简报)
- 技术背景
- 资深开发者
第二步 — 引擎看到了什么
在上述输入下哪些规则会被触发
每条规则都是一个针对输入的匹配模式,会在它影响的维度上加减一个固定值。这条想法会触发 7 条规则,没有任何一条做减法。
| 规则 | 触发原因 | 调整 |
|---|---|---|
| mon.subscription | 订阅是经过验证的持续性商业模式。 | +10,影响变现、竞争 |
| acq.community | 社群驱动增长具备防御性和高信任度。 | +7,影响分发 |
| bg.experienced | 资深开发者可以独立交付 MVP。 | +5,影响创始人匹配度、构建 |
| idea.detailed | 更长、更具体的想法通常意味着思考更充分。 | +3,影响市场 |
| val.cost_low | 验证可以通过落地页或候补名单低成本完成。 | +3,影响验证 |
| audience.icp_clear | 目标用户定义具体(明确岗位/行业/收入/待完成工作)。 | +4,影响市场、验证 |
| evidence.dense | 想法中包含具体数字、数据或试点证据。 | +3,影响市场、验证 |
第三步 — 各维度分数
基线 + 调整,夹在 0–100 之间
每个维度从基线开始。每条影响该维度的规则做加减。结果被夹在 0–100 的整数区间。
| 维度 | 基线 | 调整 | 分数 |
|---|---|---|---|
| 市场需求 | 45 | +3, +4, +3 | 55 |
| 竞争格局 | 50 | +10 | 60 |
| 分发渠道 | 40 | +7 | 47 |
| 变现模式 | 45 | +10 | 55 |
| 构建难度 | 55 | +5 | 60 |
| 创始人匹配度 | 50 | +5 | 55 |
| 验证可行性 | 50 | +3, +4, +3 | 60 |
第四步 — 两个汇总数字
总体分 vs 机会分
总体分是七个维度的加权和;机会分是同等权重的简单平均。当两者出现差距时,这个差距本身就是信号。当两者贴近时,意味着没有维度在掩盖另一个维度。
总体分(加权)
55
0.25·验证 + 0.20·分发 + 0.18·市场 + 0.12·变现 + 0.10·竞争 + 0.10·创始人匹配 + 0.05·构建 = 55.4 → 55
机会分(简单平均)
56
(55 + 60 + 47 + 55 + 60 + 55 + 60) ÷ 7 = 56.0
差距 = 1
两个数字几乎重合。没有任何维度是「引擎给低权重、但创始人很强」,也没有任何维度是「引擎给高权重、但创始人很弱」。这套输入是平衡的。
第五步 — 结论
阈值规则如何读取这些分数
规则
Go 要求总体分 ≥ 75 且没有任何单项低于 50。Caution 要求总体分 ≥ 55 且没有任何单项低于 30。其余均为 No-go。
应用到这个想法
- Go:总体分 55 低于 75 — 不满足。
- Caution:总体分 55 达到下限,且最低维度为 47(≥ 30)— 满足。
结论:Caution。
这意味着什么
输入的结构——清晰的 ICP、密集的证据、明确的渠道、订阅制变现——本身已经做了大量工作。验证分达到 60,是因为输入里包含具体的候补名单规模和访谈数量,而不是因为产品好。分发分停留在 47,是因为社群是一个可靠的首个渠道,但它本身不构成复利。 没有任何维度处于危险区(全部 ≥ 30),也没有任何维度强到能越过 Go 的门槛(全部 < 75)。现在的决定不是「立刻开做」,而是「先验证付费意愿和续费信号,再扩展构建」。
第六步 — 下一步实验
用最低成本针对剩余风险的实验
候补名单已经证明了「表达出的兴趣」。剩下的未知是这些商家是否会按订阅价付款并留下来。在不动手扩展构建的前提下,测试这件事最便宜的实验是:拿 5 位候补名单成员,按拟定价格(比如每月 99 美元)做 30 天的付费试点。 试点的任务不是验证兴趣——候补名单已经做过了。试点的任务是验证续费。没有续费信号,候补名单只是「兴趣」的证据,不是「付费」的证据。
这一结论不能证明什么
55 分的 Caution 不是对这个想法的判决,而是对「输入本身说了什么」的结构化解读。引擎没有见过这些客户,没有看过意向书,也没有测算过短信通道的成本。 同样的想法换一个获客渠道,分数会不同。把渠道从社群改成「未确定」,分发分就从 47 降到 34;总体分从 55 降到 53。这种变化可以追溯到单一项输入。 这种可追溯性,正是分数不交给语言模型决定的全部原因。
为什么重要
验证带来学习,而不是确定性
Founders validate because building is expensive and early confidence is often based on untested assumptions. A structured process helps turn a broad idea into a short list of questions that can be answered through customer behavior. Startup MRI makes that list easier to inspect, while real conversations and experiments remain the source of truth.
此方法不能做到什么
AI cannot predict startup success, replace customer research, or know facts that were never provided. Scores are decision-support signals produced from structured inputs, not probabilities of success. Treat the report as a starting point for experiments and update your view when new evidence disagrees.
Evidence re-evaluation
What evidence can change in a re-evaluation
After you run the Recommended Experiment from the report, you can submit what happened on the report page itself. Startup MRI runs the same deterministic engine against your new evidence and surfaces a side-by-side comparison: original assessment, evidence you submitted, updated assessment, what changed, and the next recommended action. The model is grounded in published customer-development and lean-startup literature (see Sources). It does not invent offsets; every change is conservative and tied to the kind of evidence you collected.
Evidence strength is classified, not counted
Evidence is classified into three tiers by kind, not by sample size alone. Weak evidence — polite interest, hypothetical answers, founder interpretation without behavioural confirmation — does not change the assessment. Moderate evidence — multiple independent descriptions of the same recent problem, observed workaround behaviour, quantified stated intent — updates the explanation. Stronger evidence — prepayment, signed LOI, quantified conversion at a non-trivial rate, completed prototype sessions — can reduce or sharpen uncertainty on the dimension the experiment targeted. More interviews alone is not stronger evidence; behavioural evidence at small N is.
Why more interviews do not automatically mean a higher score
Three reasons. First, polite interest is not commitment — The Mom Test classifies compliments as the fool's gold of customer learning. Second, quantity alone is insufficient — Strategyzer distinguishes behavioural evidence from stated evidence from opinion. Third, the original assessment is not invalidated by positive feedback on one experiment; the model requires the evidence to clear three gates before any uncertainty status flips.
What evidence re-evaluation cannot do
It cannot predict startup success. It cannot lift a dimension the experiment did not target. It cannot make Weak evidence into Stronger evidence by collecting more of it. It cannot override the original engine rules. The largest positive offset is +8; the largest negative offset is -12. A single re-evaluation cannot flip the verdict unless at least two dimensions changed OR a Stronger + Moderate pair agrees on the same dimension.
常见问题
创业验证方法常见问题
How does startup validation work?
Start with assumptions, identify the riskiest one, and run a small test that can produce a believable no as well as a yes.
What counts as evidence?
Recent behavior, specific past experiences, repeated workarounds, trial commitments, and payment or pilot decisions are stronger than hypothetical praise.
Why speak with customers before building?
Customer conversations can reveal the language, urgency, existing alternatives, and context that a founder cannot infer from an idea alone.
What is market validation?
It is testing whether a reachable group has a meaningful problem and whether the surrounding alternatives leave room for a useful offer.
Why is distribution part of validation?
A product idea is incomplete until you can describe how the intended customer will discover and adopt it.
How should I test monetization?
Ask for a concrete commitment such as a paid pilot, deposit, or purchase conversation instead of treating stated interest as willingness to pay.
Do the scores predict success?
No. They summarize structured inputs and tradeoffs to support a decision about what to test next.
Can AI predict startup success?
No. Startup outcomes depend on changing markets, execution, customer behavior, and factors a report cannot observe.
Turn your assumptions into a next step
Describe your idea and get a structured starting point for validation.
分析我的想法 →