Before a rising number counts as evidence
Vanity Metrics for Startups: Hits and Waitlists Aren't Evidence
Founders screenshot a waitlist of two thousand, or a hits graph that only goes up, and call it validation. Those numbers feel like progress. They do not tell you what to build, what to kill, or who to call next. Pick a short list of actionable metrics, write a tripwire before you measure, and still talk to the people behind the number.
Last updated: October 10, 2026
Direct answer
What is a vanity metric — and what should you measure instead?
A vanity metric is a number that makes you feel progress but does not tell you what to do next — classic examples are total hits, cumulative registered users, and waitlist size. An actionable metric is tied to a specific hypothesis, usually per-customer or per-cohort, and changes what you build or kill. For early startups, prefer a short list of actionable metrics plus a pre-written tripwire; treat big cumulative totals as PR, not proof.
Key takeaways
What to remember
- Collect metrics that help you decide. Vanity metrics feel good and give no next action. Actionable metrics — Ries’s example is a split-test where one group shows higher revenue per customer — change what you ship.
- Hits fail the tests: they are not people, not per-customer, poorly defined, and have no built-in causality. If you need a sub-metric to act, use that metric instead.
- When a vanity number rises, everyone credits whatever they were working on. When it falls, they blame someone else. Private realities diverge and prioritization becomes impossible.
- Metrics and A/B tests tell you that something changed, or which variant won. They often cannot tell you why. Web founders who treat dashboards as the whole of customer contact can optimize the wrong business model.
- Before you measure, write signal + threshold + window + action. Prefer split-tests that would falsify a belief, per-customer rates, and same-week cohort funnels. Keep as few metrics as you can audit back to people you can interview.
The two classes
Vanity metrics vs actionable metrics
Eric Ries, writing as a guest on Tim Ferriss’s blog (19 May 2009), put the test in one sentence: the only metrics entrepreneurs should collect are those that help them make decisions. Vanity metrics — billions of messages sent, the “GDP” of a platform, total hits — feel good and leave you with “now what?” Actionable metrics are usually per-customer or per-cohort, and they change a decision. The table is a sorting hat, not a scorecard. This page publishes no industry conversion rates.
| Class | Typical number | What it can decide |
|---|---|---|
| Vanity | Total hits, cumulative registered users, waitlist size, total messages sent | Almost nothing. “Now what?” has no answer. Treat it as PR. |
| Actionable | A split-test with higher revenue per customer; same-week cohort activation; per-segment paid rate | Roll out or kill a change. Rewrite a promise. Stop a sprint. Interview the people in the cell. |
Ries’s definition
Collect metrics that help you decide
Ries’s 19 May 2009 post is the definition this page uses. Off-the-shelf analytics default to vanity because those reports are easy and flattering. Actionable metrics take more work. The three habits below are his, paraphrased — not a Yibud invention and not a law you must run every week.
Split-tests that would falsify a belief
Ries’s illustration: add a feature with an A/B split. A few days later, group B has about 20% higher revenue per customer. Then you can roll the feature out, keep experimenting in that direction, and admit you learned something about those customers. His rule of thumb: if the test going the other way would not seriously doubt what you think you know about customers, try something bigger. The 20% is his example, not a target for your product.
Per-customer metrics — “metrics are people, too”
Total pageviews hide the difference between one person hitting a million times and a million people hitting once. Ries’s move is to look per new and returning customer, or per segment. A spike of new signups should not automatically change pages-per-person unless you acquired a different kind of customer. Aggregate “returning users” can also hide a churn wave from one viral week.
Funnel plus cohort, and measure the macro
For each week’s registrants, report what share later activated, used the product, or paid — the same people, followed forward. If those shares hold still across weeks, nothing important changed. If one jumps, you have a reason to look. Measure the outcome you actually care about (activation, retention, payment), not the click-through on a button. And keep as few metrics as possible. Detailed reports are for after you already know something is wrong.
Three principles Ries asked tool-builders to honor: measure what matters (few metrics), metrics are people (you can audit a number back to a person you can call), and measure the macro (not intermediate click vanity). This page does not treat any casual conversion remark in that post as a benchmark.
Why they hurt
Why vanity metrics are dangerous
On 23 December 2009 Ries returned to the topic on Startup Lessons Learned. The question is not “is a rising hit count at least directionally right?” It is what happens to a team that steers by that kind of number.
Hits fail every test that would make a metric usable
Hits were Ries’s favorite bad example. They count a technical process, not people. They are a gross total, not per-customer — one hit each from a million people is not a million hits from one person. Most people cannot say what counts (an image, a page, a script). And there is no built-in causality: a million hits this month does not say what caused them, how to get more, or whether they are equal. If each of those questions needs a different sub-metric, use the sub-metric.
Credit when up, blame when down — then nobody can prioritize
Ries had watched the same pattern in companies large and small. When the number rises, everyone attributes the rise to whatever they were working on. When it falls, they blame someone else. Over time each person lives in a private reality. Those realities diverge, and the team cannot agree what to do next. Functionally organized teams amplify it: each department grows its own story about why “foo projects always fail.”
Annotated charts show correlation. Actionable metrics put facts in the room.
A line chart with “key events” marked at the turns looks like an explanation. Ries’s point: at best it shows correlation, not causation. At worst it is a well-argued excuse. His heuristic is not “did we have a dashboard?” It is: in the last few crisis calls, product-prioritization meetings, and failure post-mortems — how much actionable data was in the room? Actionable metrics do not guarantee good decisions. They put facts in the room so intuition can be trained against reality instead of against whoever talked last.
Metrics are not the conversation
Metrics tell you that or which — not why
Steve Blank, writing on 17 December 2009, was not arguing against metrics. He was arguing that founders need three views of the customer, and that web teams who treat dashboards, A/B tests, and surveys as the whole of customer contact can optimize the wrong business model. Numbers are one view. They are not first-hand contact.
First-hand knowledge — leave the building, or at least the dashboard
Blank’s first view is personal observation: talk to potential or actual customers, watch the face, hear the voice. Metrics tell you that something is happening. A/B tests can tell you that one something is better than another. Neither tells you why. An online survey you cannot watch is a thin substitute. Collect the numbers. Do not let them replace the sitting.
A birds-eye view — and the view through other people’s eyes
The second view is a synthesized picture of customers, channels, and competitors — sites, social traces, win/loss notes, whatever you can assemble — so you can sketch the players on a whiteboard. The third is role-play: if you were the customer, why buy you instead of the incumbent? If you were the competitor, what would you do next? Each view is incomplete. First-hand is narrow. The birds-eye lacks ground truth. Role-play is a guess. Blank’s point is to use all three, not to pretend the dashboard is enough.
Share bad news. Customer needs are not a number you can optimize in isolation.
Blank’s cultural test: understanding a poor click-through, a weak retention number, or a lost sale matters more than cataloguing wins. Hoarding bad news is a firing offense in the companies he describes. An information culture is how a team gets to product/market fit without lying to itself. This page skips one student’s over-long survey as an anecdote. The teaching point is the same without it: a form is not a conversation.
You still need metrics, split-tests, and surveys. Blank’s warning is about treating them as customer interaction. The numeric picture can hide that you are polishing the wrong model. Talk to the people who generated the cell you are staring at.
How to pick
How to pick metrics — and write the tripwire first
This is Yibud’s suggested template, not a law and not a Ries or Blank checklist. Write the four fields before the first dashboard. Honor them on the end date. If you cannot name the decision the number would change, it is vanity wearing a nicer label.
- 1
Name the hypothesis the metric would falsify
Ries’s split-test test: if the result going the other way would not seriously doubt what you think you know about customers, pick a bigger question. “People from this landing page complete the core job” is a hypothesis. “The graph should go up” is not.
- 2
Prefer per-customer or per-segment rates, not cumulative totals
Cumulative registered users and lifetime waitlist size always rise if you keep the page up. A rate among last week’s registrants can fall. That is the point. Same-week cohorts beat a running total.
- 3
Follow one cohort through the funnel you actually care about
Same week’s registrants → did the core job → paid (or the three events that match your model). Measure those macro outcomes. Do not optimize the click on an intermediate button and call it learning.
- 4
Write signal + threshold + window + action before you look
Signal is the observable. Threshold is the miss / pass line you will honor. Window is calendar dates. Action is what you do on a miss — rewrite the promise, stop a sprint, interview the people in the failing cell. A kinder bar after the results is a recap, not a tripwire.
- 5
Keep the list short, and be able to call the people in the number
Ries: as few metrics as possible; audit a report back to individual people and get them on the phone when you do not know what the number means. If you cannot name five people inside a cell, you are watching a fog bank.
What you leave with
The metric tripwire (copy this)
If the sitting produced a dashboard and no dated action, you did not finish. Fill the brackets. Leave nothing as “traction,” “soon,” or “we’ll know it when we see it.”
The fill-in contract
Product: [name]. Hypothesis this number can falsify: [one belief about customers]. Metric class: [split-test / per-customer rate / same-week cohort funnel] — not a cumulative total. Signal: [e.g. share of last week’s landing-page signups who complete the core job in 7 days]. Threshold: [e.g. ≥25% of a sample you will honor]. Sample you will honor: [e.g. ≥40 people from that week]. Window: [start]–[end]. Action if miss: [rewrite the promise / stop the feature sprint / interview 5 people in the failing cell]. Audit: names of people inside the cell, not just a count.
The “≥40 people / ≥25% / interview 5” line is Yibud’s suggested template, not a published law and not a Ries or Blank rule. Rewrite the numbers before you measure if you have a reason. Do not rewrite them after. This page publishes no industry activation rate.
Example tripwire (Yibud's suggested template, not a law)
If, among ≥40 people who signed up last week from our landing page, fewer than 25% complete the core job in 7 days → we rewrite the promise the same day and interview 5 who signed up and did not finish. Do not add older waitlist names to chase 40. Do not replace this with total signups.
A waitlist of 2,000 with no activation tripwire stays vanity. A Yibud report score does not fire this tripwire.
The sketch
A cohort funnel you can draw without inventing a rate
Ries’s cohort report is a weekly table: people who registered in that week, then the share who later took each lifecycle step. The sketch below is the shape, not a benchmark. Fill your own events. This page does not publish a typical activation or purchase rate.
| Follow the same week’s people | Question the cell can answer | Vanity substitute to refuse |
|---|---|---|
| Registered this week (named source) | Did we acquire the customers we said we would, from the channel we named? | Lifetime registered users. Hits. “Traffic is up.” |
| Completed the core job (your definition, written first) | Did this week’s people do the job the landing page promised? | Button click-through. Time-on-site. “Engagement.” |
| Paid, or the scarce commitment you named | Did the same people exchange money or a scarce resource — not a compliment? | Waitlist size. Total messages sent. Cumulative accounts. |
If the shares hold still week to week, you have not learned a new thing — which is itself information. If one cell moves, you have a reason to call the people in it. Do not annotate the chart with a launch date and call the wiggle causation.
Not the same page
Vanity metrics vs NPS, sample size, Ellis, money tests, interviews, Yibud scores, and kill lines
Several Yibud pages sit next to this one. Mixing them up turns a waitlist, a recommend score, or a mid-band engine number into “the market is moving.”
NPS — one recommend-intent score, not metric hygiene
That page is stated recommendation intent among people who already used you (promoters minus detractors). This page is the general sorting hat: vanity vs actionable. A high NPS can still be vanity if it is a waitlist hypothetical or a number with no tripwire. An actionable activation rate is not an NPS.
NPS for early-stage startups →Smoke-test sample size — how large N must be, not which class of metric
That page reads k signups out of n qualified visitors as a 95% Wilson interval. This page asks whether “signups” is even the right class of number. A precise vanity metric is still vanity. A small-n actionable rate is noisy — take that problem to the sample-size page — but it is aimed at a decision.
How many visitors a smoke test needs →Sean Ellis PMF survey — disappointment among current users
That page’s signal is the percent of recent real users who would be very disappointed if the product disappeared. It is one instrument, after usage exists. This page is what to refuse before you have users: hits, lifetime signups, waitlist size as proof. Do not treat a rising dashboard as Ellis’s question.
Product-market fit survey →Willingness to pay and smoke tests — behavior and money beat a feel-good total
A deposit, pre-order, paid concierge, or checkout is money moving. A smoke-test click is reach and intent before usage. Both can be actionable if you wrote a tripwire. Neither is “2,000 emails.” Do not file a waitlist under money.
Mom Test and customer discovery — the “why” Blank says metrics miss
Blank’s first-hand view is a conversation, not a cell. The Mom Test is last-week behavior with no pitch. Customer discovery is how you find and talk to people. When a cohort cell moves and you do not know why, those pages are the next sitting. They are not a substitute for an actionable rate, and the rate is not a substitute for them.
Yibud scores — not a vanity dashboard, and not a success prediction
The validation-score and score-calculator pages explain Startup MRI numbers from a deterministic rule engine on a one-line idea plus five questions. Those scores point at the weakest assumption to test next. They are not hits. They are not a waitlist. Do not read a mid-band Yibud score as traction, and do not read a rising dashboard as a Yibud GO.
Kill criteria and critical assumptions — where a tripwire becomes a decision
Those pages write Stop, or name the one claim that would kill the idea if false. This page can supply the metric class those contracts later read: a dated, per-cohort rate against a bar you wrote first, audited back to people. A vanity total cannot fire a kill line honestly.
Worked example
One waitlist week you can copy (illustrative — fictional)
The week below is illustrative. Names, product, and counts are invented so you can see the setup. It is not a study, not an industry benchmark, and not a Yibud report. Copy the method, not the numbers.
The promise, written before the first dashboard refresh
Illustrative names: founder Priya Shah at Batchnote, a tool that turns a restaurant pre-shift meeting into a shared prep list. The landing page promises that a closer can publish tonight’s list and the opening cook will complete it before the door unlocks. The tempting trophy is a waitlist of 2,000 emails collected over three months. That total has no activation tripwire. The current bet is: people who sign up from this page this week will do the core job — publish a list that a second person completes — within 7 days.
The vanity read, and why it decides nothing
2,000 waitlist emails. Hits on the landing page are “up.” Cumulative registered users ticked again yesterday. None of those numbers say whether last week’s people did the job, whether the promise is false, or who to call. Ries’s “now what?” applies. Celebrating 2,000 is PR.
The tripwire, written the same sitting
Signal: among people who signed up last week from this landing page, the share who complete the core job in 7 days, reported with n. Threshold: honor the read only at ≥40 such people; miss if fewer than 25% finish. Window: that week’s Monday through the following Sunday, plus 7 days of follow. Action if miss: rewrite the promise the same day and interview 5 who signed up and did not finish. Older waitlist names do not enter n. A Yibud score does not fire this line.
What a miss looks like — and what must not happen
Illustrative close: 47 signups from that week. 8 completed the job in 7 days. 8/47 is under 25%, n = 47. Day 8 is a miss. The 2,000-name list is still 2,000 and still not the metric. Do not add older emails to dilute the miss. Do not switch the story to button CTR because that graph looks kinder. Do not treat the miss as “almost a good cohort.” Interview the 39. Do not invent a “typical restaurant-SaaS activation rate” this page does not publish.
Write the tripwire before the first refresh. If the window closes and the cohort misses the bar you wrote, honor the action. If it cleared the bar, continue to the next cheapest test — usually money or a named retention event, not a factory. The method is the dated page, not the waitlist screenshot.
Where Yibud fits
Neither a vanity dashboard nor a Yibud score predicts success
Yibud is a free, no-signup startup idea validator. Scores come from a deterministic rule engine, not from a language model guessing success. Optional AI text, when it is used, only polishes prose. It does not invent the number. A Startup MRI report can name a weak “people will show up” or distribution assumption, then a Day-7 Continue / Refine / Re-test / Stop signal. The score is a flashlight on the weakest assumption to test next — not traction, not hits, not a waitlist. Metric work sits next to that assumption: write the tripwire, do not hope a rising dashboard means product-market fit. If you already have a report, start with the named assumption, not the overall score. Then run the window yourself. The method on this page stands alone if you never open the analyzer.
Analyze my idea →Common mistakes
What usually wastes an early metrics week
Celebrating waitlist size
A waitlist of 2,000 with no activation tripwire is vanity. It does not say the promise works. Write a per-cohort job-completion line, or take the money question to the willingness-to-pay page. Do not screenshot the total and call it evidence.
Reporting cumulative users without cohorts
Lifetime registered users only go up if the page stays up. Ries’s move is the same week’s people, followed through activation and payment. A running total cannot falsify a belief about this week’s customers.
Optimizing click-through while ignoring retention or payment
Ries’s “measure the macro”: even when you split-test a button, the outcome that matters is the customer behavior that leads to something useful — purchase, retention, or the success your model named. A higher CTR on a button you do not care about is intermediate vanity.
Never talking to the people behind the number
Ries: metrics are people; if you do not know what the number means, get those customers on the phone. Blank: metrics and A/B tests tell you that or which, not why. A cell you cannot audit back to five names is a fog bank.
Treating a rising dashboard as product-market fit — or as a forecast
A graph that goes up is not Ellis’s very-disappointed question, not payment, and not a Yibud GO. Neither a vanity dashboard nor a Yibud report score predicts that the startup will succeed. A tripwire you dated before you looked is a decision aid. A trophy is not.
Sources
Where these ideas come from
- Eric Ries, “Vanity Metrics vs. Actionable Metrics,” guest post on Tim Ferriss, 19 May 2009 — Used for the decision test (only collect metrics that help you decide); vanity as feel-good numbers without a next action (hits, message totals); the A/B illustration (~20% higher revenue per customer in one group); split-tests, per-customer metrics (“metrics are people, too”), funnel and cohort analysis; and the three principles (measure what matters, metrics are people, measure the macro). This page does not treat any casual conversion remark in that post as a target or a benchmark.
- Eric Ries, “Why vanity metrics are dangerous,” Startup Lessons Learned, 23 December 2009 — Used for hits as the archetype vanity metric (not people, not per-customer, unclear definition, no causality); credit-when-up / blame-when-down and diverging private realities; annotated “key event” charts as correlation, not causation; the heuristic of how much actionable data was in the room for crisis, prioritization, and post-mortem decisions; and the line that actionable metrics do not guarantee good decisions but put facts in the room.
- Steve Blank, “Building a Company with Customer Data – Why Metrics Are Not Enough,” 17 December 2009 — Used for the three views (first-hand, birds-eye, through customers’ and competitors’ eyes); metrics and A/B tests as that / which, not why; the warning that web founders can confuse metrics, tests, and surveys with customer interaction and optimize the wrong model; and an information culture that shares bad news. This page does not treat one student’s survey as a universal statistic.
In one paragraph
Summary you can quote
A vanity metric is a number that feels like progress and does not tell you what to do next — total hits, cumulative registered users, waitlist size, message totals. An actionable metric is tied to a hypothesis, usually per-customer or per-cohort, and changes what you build or kill. Ries (Tim Ferriss, 19 May 2009): only collect metrics that help you decide; his split-test illustration is one group with higher revenue per customer; prefer split-tests, per-customer rates, and cohort funnels; measure few metrics, audit them back to people, and measure the macro outcome, not button vanity. Ries (23 December 2009): hits fail those tests; rising numbers get claimed by everyone, falling numbers get blamed on someone else; annotated charts show correlation; actionable metrics put facts in the room — they do not guarantee a good decision. Blank (17 December 2009): metrics and A/B tests tell you that or which, not why; founders need first-hand conversations plus a birds-eye view plus the view through customer and competitor eyes. Before you measure, write signal + threshold + window + action. Neither a vanity dashboard nor a Yibud score predicts success.
FAQ
Questions founders actually ask
What is a vanity metric?
A number that makes you feel progress but does not tell you what to do next. Ries’s classic examples are total hits, cumulative registered users, and PR totals such as messages sent. A waitlist size with no activation tripwire belongs in the same class.
What is an actionable metric?
A metric tied to a specific hypothesis, usually per-customer or per-cohort, that changes a decision. Ries’s illustration is an A/B test where one group shows higher revenue per customer — then you can roll the change out or keep learning in that direction. Actionable metrics do not guarantee a good decision. They put facts in the room.
Is a waitlist of 2,000 validation?
No. It is a cumulative total with no built-in next action. An actionable version names a cohort and a job: among people who signed up last week from the landing page, did enough of them complete the core job in 7 days, against a bar you wrote first? This page publishes no “good” waitlist size and no typical activation rate.
Are metrics enough without talking to customers?
No. Blank: metrics and A/B tests tell you that something changed, or which variant won; they often cannot tell you why. Ries: if you do not know what a number means, get the people who generated it on the phone. Dashboards without conversations can optimize the wrong business model.
Is a Yibud score a vanity metric? Does it predict success?
No and no. Yibud scores come from a deterministic rule engine on your idea inputs. They point at the weakest assumption to test next. They are not hits, not a waitlist, and not a forecast. Neither a vanity dashboard nor a Yibud report score predicts success. You do not need an account to use this page.
Write the metric tripwire before the next dashboard refresh
Yibud scores weak assumptions with a deterministic engine. You write signal, threshold, window, and action while you are still honest — then a cohort rate is a tripwire, not a waitlist trophy.
Generate a Yibud report →