Skip to content
YibudYibud

Score interpretation

What a Startup Validation Score of 60 Actually Means

Founders who get a mid-60s overall score often ask whether that is “good enough to build.” The short answer: treat it as a caution signal to design a cheap test—not as permission to ship a full product.

Last updated: September 13, 2026

Direct answer

Is 60 a good validation score?

On Startup MRI (Yibud), an overall score around 60 usually maps to CAUTION when no dimension is below 30: the idea is not rejected, but it is not a GO either. GO requires overall ≥75 and every active dimension ≥50. A lone number never predicts startup success; it only ranks relative risk across the dimensions the engine measures.

Key takeaways

Six things to remember before you act on a 60

  • Read the decision label (GO / CAUTION / NO-GO) before celebrating or panicking about the overall number.
  • 60 sits in the mid CAUTION band under published rules—if every dimension stays above the critical floor.
  • One dimension below 30 can force NO-GO even when overall still looks like “60.”
  • Overall is a weighted summary; validation and distribution carry more weight than build difficulty.
  • 60 is not a success probability, percentile, or industry benchmark.
  • The useful next move is one cheap experiment on the weakest assumption—not a larger build.

Why this matters

Why founders misread mid-band scores

School grades trained people to treat 60 as “barely pass.” Startup scores do not work that way. A mid-band overall can hide a broken channel, weak willingness to pay, or zero validation evidence. Misreading 60 as green light burns months; misreading it as total failure can discard a fixable idea. The job is to interpret the score against explicit rules, then pick the next evidence step.

Decision bands

How Startup MRI turns scores into GO / CAUTION / NO-GO

These thresholds are product rules published in the methodology and implemented in the scoring engine. They are not third-party industry standards.

DecisionOverall scoreDimension guardrails
GO≥ 75Every active dimension ≥ 50 (no “weak” dim)
CAUTION≥ 55 (and below GO)No dimension below 30 (no “critical” dim)
NO-GOBelow 55, or fails a guardrailAny critical dim (<30), or overall too low even if dims look fine

Source: Yibud decision thresholds (DECISION_GO_MIN_OVERALL = 75, DECISION_CAUTION_MIN_OVERALL = 55, weak floor 50, critical floor 30). Full write-up lives on the Methodology page.

Reading a 60

What “about 60” usually means in practice

If overall is ~60 and the weakest dimension is still ≥30, the engine’s decision is CAUTION. That means: keep investigating with low commitment; do not treat the score as a build mandate. If overall is ~60 but any dimension is under 30, the decision flips to NO-GO—the average hid a fatal gap.

When ~60 is actionable caution

Overall mid-50s to mid-60s, no critical dimension, and you can name the weakest assumption. Next step: one experiment that can falsify that assumption cheaply.

When ~60 is still a stop

A critical dimension below 30, or you plan to “fix the score” by rewriting inputs without new evidence. The number is then noise; the broken dimension is the signal.

How to read the report

A five-step reading order

Use this sequence every time you open a report—especially around 60.

  1. 1

    Check the decision label first

    GO, CAUTION, or NO-GO already encodes overall plus dimension floors. Do not invent your own pass/fail from the overall alone.

  2. 2

    Find the lowest dimension

    Market, distribution, monetization, competition, build, founder fit, and validation each tell a different risk story. The lowest one is usually where the next test belongs.

  3. 3

    Separate “hard to build” from “hard to sell”

    A weak build score suggests scope or skill risk. A weak distribution or monetization score suggests demand and channel risk. Those need different experiments.

  4. 4

    Treat overall as a weighted summary

    Raising a low-weight dimension barely moves overall; raising validation or distribution moves it more. Chasing vanity points on easy dimensions is a common trap.

  5. 5

    Write one falsifiable next step

    Turn the weakest claim into a test you can run this week—interview, fake-door, concierge delivery, or a pre-order style commitment—then re-evaluate with evidence.

Weights (product rules)

How overall is composed

Current V4 active-dimension weights used by the rule engine. These explain why two reports with the same “feel” can score differently.

  • Validation25%
  • Distribution20%
  • Market18%
  • Monetization12%
  • Competition10%
  • Founder fit10%
  • Build5%

Source: published WEIGHTS in the Startup MRI scoring rules (methodology). Weights can change across engine versions; always trust the report version you just generated.

Common misreads

What a score of 60 does not mean

Not a 60% chance of success

The engine does not estimate probability of fundraising, product-market fit, or revenue. Claiming otherwise would invent a forecast the product does not make.

Not a percentile vs other startups

Scores are absolute against rule tables for your inputs—not ranked against a public dataset of other founders’ ideas.

Not permission to hire or raise

CAUTION still means high uncertainty. Use it to justify cheap learning, not payroll or a large build.

Not proof that demand exists

Inputs can be optimistic. Without interviews, commitments, or traffic tests, a mid score can still rest on untested assumptions.

What to do next

If your overall is around 60

Prefer evidence over more analysis theater.

  1. 1

    Name the critical assumption

    Write one sentence: “We believe [customer] will [pay / switch / use weekly] because [reason].” If you cannot write it, you are not ready to interpret the score.

  2. 2

    Pick one cheap test

    Match the weak dimension: distribution → landing or fake-door; monetization → willingness-to-pay conversations; validation → structured customer discovery.

  3. 3

    Define kill/continue criteria before you run it

    Decide what evidence would make you stop, pivot the wedge, or keep going. Changing the bar after results is how mid scores become self-justifying.

  4. 4

    Re-evaluate with new evidence

    Update inputs honestly after the test. The useful outcome is a clearer decision—not a higher vanity number.

Failure modes

Five mistakes founders make with mid-band scores

Averaging away a critical dim

Celebrating overall 62 while distribution sits at 22. The guardrail exists because averages lie.

Re-running the form until the number rises

Changing adjectives without new facts. That optimizes the report, not the business.

Treating GPT prose as stronger than scores

On Startup MRI, scores come from deterministic rules; narrative explains them. Do not let flattering language override a weak dimension.

Jumping from CAUTION to a six-month build

Mid-band means uncertainty is still high. Scale commitment only after evidence moves the decision.

Ignoring validation weight

Validation is the heaviest dimension. Skipping interviews or demand tests while polishing features fights the scoring design.

Limits

When this page (and the score) should not decide for you

Do not use a single report as legal, investment, or employment advice. The engine cannot see offline relationships, regulatory blockers, or team execution quality beyond what you entered. If your market is highly regulated or safety-critical, treat the score as a conversation starter with domain experts—not a clearance.

Sources

Authoritative references used on this page

  • Yibud / Startup MRI MethodologyPrimary source for decision thresholds, dimension floors, and V4 weights described above.
  • The Lean Startup — Principles (Eric Ries)Frames “validated learning”: progress is measured by tested assumptions, not by polishing an untested plan. Used here to justify experiments after a mid-band score—not to invent success rates.
  • Steve Blank — Customer DevelopmentSupports getting out of the building when demand and channel assumptions are unproven. Used for next-step framing, not numerical benchmarks.

FAQ

Questions founders ask about a score of 60

Is a validation score of 60 good?

It is usually actionable caution, not a green light. Under Startup MRI rules, ~60 typically lands in CAUTION if no dimension is critically low. Good enough to keep learning; not good enough to treat as GO.

Can I get GO with a 60 overall?

No. GO requires overall ≥75 and every active dimension ≥50. Raising overall into the 70s while leaving a dim below 50 still blocks GO.

Why did I get NO-GO with overall near 60?

Likely a critical dimension below 30, or overall below the CAUTION floor after guardrails. Open the breakdown and find the dim under 30 first.

Does 60 mean my idea is average?

No. The score is not a percentile against other startups. It is a weighted rule output for your inputs.

Should I rebuild the product until the score hits 75?

No. Change inputs only when you have new evidence. Aim for clearer risk, not a vanity climb.

How is this different from ChatGPT giving me a score?

Startup MRI scores are computed by deterministic rules; the model explains them. Chat-only scores can invent thresholds. Prefer published rules you can inspect on the Methodology page.

Want your own breakdown—not just a headline number?

Run Startup MRI, then read the decision label and weakest dimension with this page open. Use the score to pick one experiment, not to declare victory.

Analyze my idea