Validation Guide
How to Size a Market You Can Actually Reach: TAM, SAM, SOM, and the Two Numbers That Are Usually Useless
A founder's guide to estimating the market they can actually reach — what TAM, SAM, and SOM actually mean, why two of the three are usually useless for solo founders, how to do a defensible bottom-up sizing in a single afternoon, where to source the data, and the failure modes that produce a market-size claim a partner or investor will not trust.
· Updated · Yibud· 16 min read
On this page
A founder pitches a small CRM to an angel investor. The deck claims the total addressable market is $61 billion. The investor does not blink. The deck claims the serviceable addressable market is $24 billion. The investor does not blink. The deck claims the serviceable obtainable market is $900 million. The investor asks, quietly, how the founder got the number. The founder does not remember. The deck was copied from a McKinsey template. The number was rounded up.
The number was a lie. Not an intentional one — the founder had used the most common shortcut in startup market sizing, which is to take an industry analyst's top-line figure, apply a percentage, and round. The shortcut produces a number that satisfies the audience of the pitch deck and reveals nothing about whether the founder's specific product can serve the specific people the founder has decided to serve. The number is the kind of claim that lives in pitch decks because the deck format requires it, not because anyone believes it.
This article is the alternative. It is the founder's guide to estimating the market they can actually reach, with numbers that survive contact with a partner who asks how they were produced. The article draws on Bill Aulet's Disciplined Entrepreneurship (MIT, 2013), which is the canonical academic treatment of the methodology for solo founders; Alexander Osterwalder and Yves Pigneur's Business Model Generation (2010), which is the source of the Value Proposition Canvas; the McKinsey Valuation textbook's bottom-up-vs-top-down sizing chapter; and the practitioner tradition that lives in Paul Graham's essays, the Y Combinator curriculum, and the public market-sizing work published by Andreessen Horowitz (a16z) and First Round Capital.
The piece has four parts. The first names what TAM, SAM, and SOM actually are, and the two of them that are useless for a founder who is not raising venture capital. The second describes the only sizing exercise that produces a number worth defending: bottom-up sizing from the named buyer. The third names the cheapest data sources, with primary-source links, that a founder can pull in an afternoon. The fourth names the failure modes that produce a market-size claim a partner or investor will spot in a minute.
If you have ever been told your market size is "too small" without anyone explaining how a market size could be too small for a business, this is the article that explains why the answer is rarely the number — it is the sizing methodology behind the number.
Quick answer
Market sizing for a startup is the discipline of estimating three nested numbers: the total addressable market (TAM) — every possible buyer of the category worldwide; the serviceable addressable market (SAM) — the slice of the TAM the founder's product and geography can serve; and the serviceable obtainable market (SOM) — the slice of the SAM the founder can realistically reach in three years given the current channel, ICP, and team.
For a solo founder, TAM and SAM are usually useless. TAM is the kind of number a McKinsey report cites for an industry press release; SAM is the kind of number that satisfies a pitch-deck template. Neither produces a decision the founder can make tomorrow. SOM is the only market-size number that matters for a solo founder, because SOM is the number that connects directly to revenue, distribution, and the channel the founder will actually use.
The sizing exercise that produces a defensible SOM is bottom-up sizing from the named buyer: start with the number of people in the named ICP, multiply by the per-customer annual price, and adjust by the conversion rate the channel will realistically produce. The result is a number with traceable inputs. The exercise takes a single afternoon if the founder already has an ICP; longer if the ICP itself has to be defined first.
The data sources a founder should use are first-party government statistics (US Census Bureau, BLS, Eurostat), industry-association reports with named methodology, the founder's own comparable-set data, and channel-driven estimates from the platforms the founder will actually use (Google Keyword Planner, Reddit subscriber counts, Slack community directories, Crunchbase for competitor revenue ranges). The data sources a founder should avoid are rounded analyst forecasts without a named methodology, second-hand "the market is $X" claims copied between pitch decks, and synthetic per-customer math that multiplies a category size by a percentage with no traceable basis.
Key takeaways
- TAM is the largest possible market. SAM is the slice the founder's product can serve. SOM is the slice the founder can actually reach in three years. The three numbers are nested, not independent.
- For a solo founder, the SOM is the only market-size number that matters. TAM and SAM are the numbers a pitch deck template requires; SOM is the number that connects to revenue, distribution, and the channel the founder has actually chosen.
- The sizing exercise that produces a defensible SOM is bottom-up. Start with the named ICP, multiply by the annual price, and adjust by the channel conversion rate. The inputs are traceable. The number can be defended.
- Top-down sizing produces numbers that are easy to compute and impossible to defend. "X% of a $Y billion market" is the formula. The output is a rounded number no one will challenge in a deck and no one will remember afterwards.
- First-party data sources (US Census, BLS, OECD, Eurostat) are the most defensible. Industry-association reports are second. Comparable-set data from competitors is third. Analyst forecasts are last and weakest.
- A market-size claim with no named methodology is a market-size claim with no value. A partner or investor who asks how the number was produced and receives a "well, it's roughly X% of the global market" answer has just heard the founder admit the number was rounded.
- A market-size claim that survives scrutiny is a market-size claim with three named inputs. Number of buyers in the named ICP, per-customer annual price, channel conversion rate. Every input has a source. Every source is fetchable.
Why this matters
Market sizing is the part of a startup's evidence base that founders most often fake, and the part that an investor or partner is most able to spot. The reason is that most market-size claims are produced by one of three shortcuts: copying an analyst headline ("the global CRM market is $61 billion"), multiplying a category size by an arbitrary percentage ("our SOM is 1% of $61 billion"), or counting the addressable population as buyers ("there are 4 million independent insurance agencies in the US, so our TAM is 4 million"). Each shortcut produces a number that satisfies the format of a pitch deck and reveals nothing about whether the founder's specific product can serve the specific buyers the founder has decided to serve.
The cost of the shortcut is not theoretical. A founder who raises against a $900 million SOM and finds, eighteen months later, that the reachable market is closer to $9 million has not failed at distribution — they failed at sizing. The capital the founder raised was committed against a number that did not exist. The time the founder spent was spent chasing a market the founder did not have. The discipline of market sizing, done honestly, is the discipline of distinguishing the market the founder wants to address from the market the founder can actually reach.
The second reason sizing matters is that the question founders most often answer wrong is not "is there a market?" but "is the market the founder can reach big enough to be worth building for?" A market the founder can reach at $50,000 in annual recurring revenue is not a market the founder should commit three years of work to. A market the founder can reach at $5 million in annual recurring revenue is. The number that distinguishes the two is the SOM — the slice of the SAM the founder can realistically convert in three years with the channel the founder has chosen. A founder who has not done this calculation is guessing about whether the idea is worth three years of work.
Definitions
- TAM (Total Addressable Market) — the total annual revenue the entire category could produce if 100% of the possible buyers bought the category's products. TAM is a category-level number; it has nothing to do with the founder's product, geography, or channel. For a CRM, the TAM is the total revenue the global CRM category could produce if every possible CRM buyer bought a CRM.
- SAM (Serviceable Addressable Market) — the slice of the TAM the founder's product, geography, business model, and regulatory environment can serve. For a US-based CRM focused on independent insurance agencies, the SAM is the slice of the global CRM TAM that US-based independent insurance agencies represent. The SAM excludes buyers the founder cannot serve (because of geography, regulation, language, or product scope).
- SOM (Serviceable Obtainable Market) — the slice of the SAM the founder can realistically reach in three years given the channel, ICP, and team the founder has today. The SOM is the slice of the SAM the founder will actually convert into customers. The SOM is the only market-size number that connects to revenue the founder can plan against.
- Bottom-up sizing — the methodology that starts with the named ICP, multiplies by the per-customer annual price, and adjusts by the channel conversion rate. The output is a number with traceable inputs. Bottom-up sizing is the methodology that produces a defensible SOM.
- Top-down sizing — the methodology that starts with an analyst's headline figure for the category and applies a percentage ("X% of a $Y billion market"). Top-down sizing produces numbers that are easy to compute and impossible to defend. Top-down sizing is the methodology that produces the "round the number up" lie.
- ICP (Ideal Customer Profile) — the narrow description of the one person whose problem the founder's product solves, in enough detail that the founder can name a room they actually gather in. The narrower the ICP, the more accurate the SOM.
- Channel conversion rate — the fraction of people the founder reaches through the named channel who become customers. The channel conversion rate is the input that converts a population count into a customer count.
The two numbers that are usually useless
TAM and SAM are the two market-size numbers that are usually useless for a solo founder. Each fails for a specific reason.
Why TAM is usually useless
TAM is the total annual revenue the entire category could produce if 100% of the possible buyers bought the category's products. For a CRM, TAM is the total global revenue every CRM in the world could produce if every possible CRM buyer bought one. The number is enormous, the number is rounded, and the number has nothing to do with the founder's specific product, ICP, or channel.
The reason TAM is useless for a solo founder is that no solo founder will ever serve the entire TAM. The TAM is the kind of number that an analyst at McKinsey or Gartner produces for an industry press release. The TAM is also the kind of number that gets copied between pitch decks for a decade, with each copy rounded up a little further, until the number has no traceable source and no relationship to anything the founder has actually measured. A founder who quotes a TAM in a pitch deck has not done sizing. The founder has cited a headline.
There is one situation in which TAM is worth quoting: when the founder is raising venture capital at a scale where the venture return requires a category-level outcome. A founder building a CRM who is raising $50 million to acquire 1% of the global CRM market is making a TAM-shaped bet, and the TAM is the relevant scale of the bet. A solo founder raising $50,000 to acquire 0.0001% of the global CRM market is not making a TAM-shaped bet, and quoting the TAM in the deck is the lie.
Why SAM is usually useless
SAM is the slice of the TAM the founder's product, geography, business model, and regulatory environment can serve. SAM is more specific than TAM and more useful than TAM, but still too large to drive a solo founder's decisions.
The reason SAM is usually useless for a solo founder is that SAM is still a slice of the category, not a slice of the founder's actual reach. A founder's SAM might be $24 billion (the revenue US-based independent insurance agencies could spend on a CRM if every one of them bought a CRM). The SAM is the right shape of number — it accounts for the founder's geography, the founder's product scope, and the founder's regulatory environment. The SAM is still a number no solo founder will ever serve. The SAM is the category the founder has decided to play in, scaled down to the slice the founder could play in. The founder is not playing in the slice; the founder is playing in the SOM.
Why SOM is the only number that matters
SOM is the slice of the SAM the founder can realistically reach in three years given the channel, ICP, and team the founder has today. SOM is the only market-size number that connects directly to revenue, distribution, and the channel the founder has actually chosen.
The reason SOM is the only number that matters is that SOM is the answer to the question "is the market I can actually reach big enough to be worth three years of work?" A market the founder can reach at $50,000 in annual recurring revenue is not. A market the founder can reach at $5 million is. The number that distinguishes the two is the SOM. A founder who has not calculated the SOM has not answered the question that decides whether the idea is worth committing to.
SOM is also the only market-size number that has traceable inputs. SOM is the product of three named numbers: the number of buyers in the named ICP, the per-customer annual price, and the channel conversion rate. Each input has a source. Each source is fetchable. The methodology is reproducible. The output can be defended.
The bottom-up sizing exercise
The sizing exercise that produces a defensible SOM is bottom-up sizing from the named buyer. The exercise has six steps, each of which has a named input and a named source.
Step 1 — Name the ICP
The first step is to name the ICP in enough detail that the founder can count the population. "Independent insurance agencies in the US" is too broad. "Independent insurance agencies in the US with 10 to 50 agents that currently use a spreadsheet to manage renewals" is specific enough to count.
The discipline of naming the ICP is the discipline of converting a category ("insurance agencies") into a population the founder can enumerate ("insurance agencies with 10 to 50 agents in specific US states that have a specific problem the founder's product solves"). The narrower the ICP, the more accurate the population count.
Step 2 — Count the population
The second step is to count the population of the ICP. The cheapest sources for population counts are:
- US Census Bureau (census.gov) — counts of US businesses by industry, size, and geography. The County Business Patterns and the Statistics of US Businesses datasets are the most useful.
- Bureau of Labor Statistics (bls.gov) — counts of US establishments by industry and size, with employment data.
- Eurostat (ec.europa.eu/eurostat) — counts of European businesses by industry and size.
- OECD (oecd.org) — international business demographics.
- Industry associations — most industries have an association that publishes a membership count or directory. The Independent Insurance Agents & Brokers of America, for example, publishes a member count.
For the ICP "independent insurance agencies with 10 to 50 agents in the US," the US Census Bureau's County Business Patterns dataset would be the primary source. The count would be derived by filtering the dataset by NAICS code (524210 for agencies), employment size (10–50), and corporate structure (non-franchise).
Step 3 — Apply the per-customer annual price
The third step is to multiply the population by the per-customer annual price. The per-customer annual price is the price the founder has decided to charge, expressed as annual recurring revenue per customer.
The discipline of choosing the per-customer annual price is the discipline of choosing the price before the sizing, not after. A founder who sizes the market at $9 per month and then discovers the market is too small has not done sizing; the founder has done wishful arithmetic. The price should be the price the founder has committed to, validated through a willingness-to-pay test (the discipline is in How to Test Willingness to Pay Before You Build and How to Price a New SaaS Product).
Step 4 — Apply the channel conversion rate
The fourth step is to multiply the result by the channel conversion rate. The channel conversion rate is the fraction of people the founder reaches through the named channel who become customers.
The channel conversion rate is the input that converts a population count into a customer count. A channel with a 2% conversion rate produces twenty times fewer customers than a channel with a 40% conversion rate, all else equal. The conversion rate is grounded in the founder's comparable-set data, not in industry averages. The discipline is named in Distribution Channels Ranked for Solo Founders.
For a solo founder using SEO, the channel conversion rate might be 1% to 3% of qualified visitors. For a solo founder using cold email, the channel conversion rate might be 0.5% to 2% of cold outreach messages. For a solo founder using a Slack community, the channel conversion rate might be 5% to 15% of community members. The rates are not averages — they are grounded in the comparable-set data the founder can pull from public case studies, comparable-set revenue disclosures, and the founder's own experiments.
Step 5 — Apply the three-year reach cap
The fifth step is to apply the three-year reach cap. The three-year reach cap is the constraint that the founder can only realistically convert a fraction of the population in three years, no matter how high the conversion rate is.
The three-year reach cap exists because distribution takes time. A founder who has chosen a channel, who has built the distribution, and who has run the experiments for three years might reach 5% to 20% of the named ICP. A founder who is just starting out will reach much less. The cap is the discipline of distinguishing the market the founder could reach in theory from the market the founder can reach in the time horizon of a serious commitment.
Step 6 — Add the traceable inputs back
The sixth step is to write down the traceable inputs and the traceable source for each input. The output of the bottom-up sizing exercise is a table with five rows:
- ICP population count, with the source.
- Per-customer annual price, with the source.
- Channel conversion rate, with the source.
- Three-year reach cap, with the source.
- SOM in annual recurring revenue, with the arithmetic shown.
The output is a table the founder can show to a partner or investor and say "this is how I sized the market, and here is the source for every input." The output is also the table the founder can defend in a one-hour conversation. The output is the discipline of bottom-up sizing.
A worked example
The example below walks through the bottom-up sizing exercise for a fictional founder building a CRM for US-based independent insurance agencies with 10 to 50 agents.
Step 1 — ICP. US-based independent insurance agencies with 10 to 50 agents that currently use a spreadsheet or a generic CRM (not a vertical-specific CRM) to manage renewals.
Step 2 — Population. The US Census Bureau's County Business Patterns dataset, filtered for NAICS 524210 with employment 10–50, returns approximately 14,000 establishments. Source: census.gov/data/datasets/cbp.html.
Step 3 — Price. The per-customer annual price is $2,400 per year ($200 per month per agency). Source: the founder's own willingness-to-pay test, which surfaced this price as the indifference point in a Van Westendorp survey of 30 prospects (the methodology is in How to Price a New SaaS Product).
Step 4 — Conversion rate. The channel is SEO + Slack community (the Independent Insurance Agents & Brokers of America Slack, plus a few adjacent communities). The comparable-set data from three vertical-CRM competitors' published case studies suggests a 4% to 8% conversion rate from qualified visitors to trial, and a 30% to 50% trial-to-paid rate, for a blended 1.5% to 4% visitor-to-paid conversion rate. The founder uses 2% as the conservative input.
Step 5 — Three-year reach cap. The three-year reach cap is 10% of the population — the founder's comparable-set data suggests that the top-three vertical CRMs in this category have cumulatively reached roughly 8% of the named ICP after five years. A solo founder starting today can reasonably expect to reach 5% to 10% in three years.
Step 6 — SOM calculation.
- ICP population: 14,000 agencies.
- Reachable in three years: 10% × 14,000 = 1,400 agencies.
- Per-customer annual price: $2,400.
- Channel conversion rate already applied: not needed if the reachable count is the founder's expected customer count, not the population count.
- SOM in ARR: 1,400 × $2,400 = $3,360,000.
The SOM is $3.36 million in annual recurring revenue. The number is small. The number is defensible. Every input has a source. Every source is fetchable. A partner or investor who asks how the number was produced can be walked through the table in fifteen minutes.
The number is also a number the founder can act on. $3.36 million in annual recurring revenue, at the founder's chosen margin, is a viable lifestyle business or a viable seed-stage raise. The founder knows whether the number is worth three years of work because the number is grounded in the founder's own comparable-set data and the founder's own channel.
Where to source the data
The cheapest primary sources for market-sizing data are below, with the use case for each.
US Census Bureau
The Census Bureau publishes four datasets that are useful for market sizing: the County Business Patterns (counts of US establishments by industry and employment size), the Statistics of US Businesses (counts by industry and firm size), the American Community Survey (population and demographic data), and the Economic Census (industry-level revenue and employment). All four are free and downloadable from census.gov/data.
The dataset the founder should reach for first is the County Business Patterns. The CBP returns establishment counts by NAICS code, employment size, and geography, with a few years of history. For most solo-founder sizing exercises, the CBP is the answer.
Bureau of Labor Statistics
The BLS publishes industry-level employment, wages, and establishment counts. The Quarterly Census of Employment and Wages is the most useful dataset for market sizing: it returns establishment counts by industry and size for every state and every county in the US. The dataset is free and downloadable from bls.gov/cew.
Eurostat and OECD
For international sizing, Eurostat (ec.europa.eu/eurostat) publishes European business demographics, and the OECD (oecd.org) publishes international business demographics. Both are free and downloadable.
Industry associations
Most industries have a trade association that publishes a member directory, a market report, or an annual statistical abstract. The Independent Insurance Agents & Brokers of America, the American Medical Association, the American Bar Association, the National Association of Realtors, and the Society for Human Resource Management all publish industry data free of charge.
Comparable-set revenue
The best source for channel-conversion-rate data is the comparable-set revenue of competitors. For a vertical CRM, the comparable-set revenue is publicly available via Crunchbase, PitchBook, and the public filings of any competitor that is public. The comparable-set revenue divided by the number of customers the competitor has disclosed gives the founder a defensible per-customer revenue figure, and the comparable-set revenue divided by the founder's estimate of the population in the ICP gives the founder a defensible penetration rate.
Platform data
For channel-conversion-rate data, the founder's own platform data is the most defensible source. Google Keyword Planner returns search-volume data for the named ICP's queries. Reddit subscriber counts return the size of the named community. Slack and Discord directories return the size of named professional communities. The platform data is what the founder will actually use, and the platform data is what the founder can defend.
Common mistakes
Seven failure modes produce a market-size claim a partner or investor will spot in a minute. Each is named below with the failure mode, an example, and the fix.
Mistake 1 — Quoting a headline as a TAM
The failure mode: a McKinsey or Gartner report cites a $61 billion TAM for the global CRM category. The founder copies the number into the deck as the TAM for the founder's CRM. The headline is for the category; the TAM the founder cites is for a specific product serving a specific ICP in a specific geography. The headline is the wrong shape.
The fix: cite the headline as the category-size backdrop, then derive the TAM by counting the population the founder's product could serve. The TAM is the founder's number, not McKinsey's.
Mistake 2 — Multiplying the TAM by a percentage
The failure mode: "our SOM is 1% of the global CRM market." The 1% is arbitrary. The founder has no basis for choosing 1% over 0.1% or 5%. The output is a number that satisfies the format of a pitch deck and reveals nothing about the founder's actual reach.
The fix: replace the percentage with the founder's three-year reach cap and the founder's channel conversion rate. The inputs are traceable. The output is defensible.
Mistake 3 — Counting the addressable population as the buyer count
The failure mode: "there are 4 million independent insurance agencies in the US, so our TAM is 4 million agencies." The 4 million is the population count, not the buyer count. The buyer count is the fraction of the population that will actually buy the founder's product, at the founder's price, through the founder's channel, in the founder's time horizon.
The fix: apply the channel conversion rate and the three-year reach cap to the population count. The output is the buyer count, not the population count.
Mistake 4 — Using an analyst forecast with no methodology
The failure mode: a deck cites a "$900 million market by 2027" figure from an analyst forecast. The methodology is not published. The inputs are not named. The forecast is the analyst's guess at what the market will be in three years. The forecast has no bearing on the founder's specific product.
The fix: avoid analyst forecasts without named methodology. Use the founder's own bottom-up sizing instead.
Mistake 5 — Sizing the market after the product
The failure mode: the founder has built the product and is now sizing the market. The sizing is biased by the need to justify the product. The founder rounds up. The founder chooses the inputs that produce the largest number.
The fix: size the market before the product. The sizing is the input to the product decision, not the output of it.
Mistake 6 — Confusing the market size with the revenue opportunity
The failure mode: the founder confuses the market size (the total annual revenue the category could produce) with the revenue opportunity (the annual revenue the founder could capture). The market size is the backdrop. The revenue opportunity is the SOM multiplied by the founder's per-customer price.
The fix: name the market size as the backdrop. Name the revenue opportunity as the SOM. The two are different numbers, and the deck should make the difference explicit.
Mistake 7 — Citing a market size without a source
The failure mode: a deck cites a market size without naming the source. The number is a rounded number with no traceability. A partner or investor who asks how the number was produced hears a vague answer.
The fix: every market-size claim in a deck cites a named source. Every source is fetchable. Every input to the sizing is in a table the founder can walk through in a one-hour conversation.
How this connects to the rest of the validation work
Market sizing is the rung before distribution. A founder who has named the ICP, validated the problem, and validated the willingness to pay can size the market and decide whether the market is worth three years of work. The output of the sizing exercise is the answer to the question "is the SOM big enough to commit to?" A "no" answer is a valid outcome. A "yes" answer is the input to the distribution work — the channel, the cadence, the budget, the team.
Market sizing also feeds the Startup MRI report's opportunity score. The opportunity score is the dimension that rewards a large, reachable, under-served market. A founder who has done the bottom-up sizing exercise has the inputs the opportunity score needs; a founder who has copied a headline has inputs that produce an inflated opportunity score and an inflated overall score. The sizing exercise is the discipline of grounding the opportunity score in the founder's own comparable-set data.
Frequently asked questions
What is the difference between TAM, SAM, and SOM?
TAM is the total annual revenue the entire category could produce if every possible buyer bought the category's products. SAM is the slice of the TAM the founder's product, geography, and regulatory environment can serve. SOM is the slice of the SAM the founder can realistically reach in three years with the channel, ICP, and team the founder has today. The three are nested, not independent.
Which market-size number matters for a solo founder?
SOM. The SOM is the only market-size number that connects directly to revenue the founder can plan against. TAM and SAM are the numbers a pitch deck template requires; SOM is the number that answers the question "is the market big enough to commit three years of work?"
How does a founder do a bottom-up market sizing?
Six steps: name the ICP, count the population, apply the per-customer annual price, apply the channel conversion rate, apply the three-year reach cap, and write down the traceable inputs. The output is a table with five rows, each citing a named source. The output is defensible in a one-hour conversation.
What is the cheapest primary source for market-sizing data?
The US Census Bureau's County Business Patterns dataset, for US-based sizing. The BLS Quarterly Census of Employment and Wages, for US establishment counts by size. Eurostat, for European business demographics. OECD, for international business demographics. Industry associations, for industry-specific membership counts. All four are free and downloadable.
Why is top-down market sizing unreliable?
Top-down sizing produces numbers that are easy to compute and impossible to defend. The shortcut — "X% of a $Y billion market" — has no traceable basis for the percentage, and the resulting number has no relationship to the founder's actual reach. The shortcut is the most common source of the "round the number up" lie in pitch decks.
How does market sizing relate to a startup's TAM claim?
For a solo founder, the TAM claim is the backdrop, not the answer. The TAM is the category the founder has decided to play in. The SAM is the slice the founder could serve. The SOM is the slice the founder can reach in three years. The number that goes into the opportunity score is the SOM, not the TAM.
Should a founder size the market before or after building the product?
Before. The sizing is the input to the product decision, not the output of it. A founder who sizes the market after the product is biased toward rounding up. A founder who sizes the market before the product has the answer the product decision needs.
How does a founder avoid the "round the number up" lie?
Every market-size claim cites a named source. Every source is fetchable. Every input to the sizing is in a table the founder can walk through in a one-hour conversation. The methodology is reproducible. The output can be defended.
Summary
Market sizing for a solo founder is the discipline of estimating the slice of the market the founder can actually reach in three years, not the slice of the category the analyst headline quotes. The three numbers are TAM, SAM, and SOM; for a solo founder, SOM is the only one that connects to a decision the founder can make tomorrow. The sizing exercise that produces a defensible SOM is bottom-up sizing from the named buyer, with six named steps and a table of traceable inputs. The data sources a founder should reach for first are the US Census Bureau, the BLS, Eurostat, the OECD, industry associations, and the founder's own comparable-set data. The data sources a founder should avoid are rounded analyst headlines, second-hand copied numbers, and synthetic per-customer math with no traceable basis.
The cost of the shortcut is the cost of committing three years of work to a market that does not exist at the size the founder assumed. The discipline of bottom-up sizing is the discipline of distinguishing the market the founder wants to address from the market the founder can actually reach.
Next action
If you have an idea and have not yet sized the market, the next concrete action is to name the ICP in a single sentence, count the population from the US Census Bureau's County Business Patterns dataset (or the BLS QCEW, or Eurostat), and write down the SOM as a single number with five traceable inputs. The exercise takes an afternoon if the ICP is already named; it takes longer if the ICP itself has to be defined first.
If you would like the market size applied to your specific ICP and idea alongside the rest of the validation dimensions, Startup MRI's validation analysis returns the SOM, the channel conversion rate, and the opportunity score in under five minutes. The analysis does not decide whether your market is worth three years of work; it organizes the inputs so the next decision is easier to make.
Sources
- Aulet, B. (2013). Disciplined Entrepreneurship: 24 Steps to a Successful Startup. MIT Press. The six-step bottom-up sizing exercise the article describes is a synthesis of Aulet's "end-user profile" step (#5) and "quantify the value proposition" step (#7). Reference at mitpress.mit.edu.
- Osterwalder, A., & Pigneur, Y. (2010). Business Model Generation: A Handbook for Visionaries, Game Changers, and Challengers. The Value Proposition Canvas and the segment-by-segment sizing logic. Reference at strategyzer.com.
- Osterwalder, A., Pigneur, Y., Bernada, G., & Smith, A. (2014). Value Proposition Design: How to Create Products and Services Customers Want. The customer-segment-by-customer-segment sizing the article uses as the input to the SOM. Reference at strategyzer.com.
- McKinsey & Company. Valuation: Measuring and Managing the Value of Companies (various editions, Wiley). The bottom-up-vs-top-down sizing chapter the article inherits. Reference at wiley.com.
- US Census Bureau. County Business Patterns and Statistics of US Businesses datasets. The free, primary-source establishment counts the article uses for US-based population sizing. Reference at census.gov/data/datasets/cbp.html.
- US Bureau of Labor Statistics. Quarterly Census of Employment and Wages. The free, primary-source establishment counts by industry and employment size. Reference at bls.gov/cew.
- Eurostat. Business Demography Statistics. The free, primary-source European business demographics. Reference at ec.europa.eu/eurostat.
- OECD. Structural Business Statistics. The international business demographics for OECD member states. Reference at oecd.org.
- Andreessen Horowitz (a16z). Market Sizing 101 (published in the a16z blog and the Fwd Startup School curriculum). The category-level-vs-segment-level sizing distinction the article uses for the TAM-is-backdrop framing. Reference summary at a16z.com.
- First Round Capital. The Startup Playbook (State of Startups survey work). The comparable-set data approach the article recommends for channel-conversion-rate inputs. Reference at firstround.com.
- Graham, P. How to Get Startup Ideas and Do Things That Don't Scale (essays). The case for sizing the SOM from the founder's own reach, not from the analyst's headline. Reference at paulgraham.com and paulgraham.com/ds.html.
- Fitzpatrick, R. (2013). The Mom Test: How to Talk to Customers & Learn If Your Business is a Good Idea When Everyone is Lying to You. The polite-yes problem that the ICP definition must clear: an ICP sized by what customers say in interviews is an ICP sized by politeness, not by behavior. Reference at momtestbook.com.
- Andreessen, M. (2007, May). Pmarca: Product-Market Fit. The PMF framing; market sizing is the rung before PMF, and the SOM is the number that defines what "big enough to commit to" means. Reference at pmarca.com.
- Ries, E. (2011). The Lean Startup: How Today's Entrepreneurs Use Continuous Innovation to Create Radically Successful Businesses. The Build-Measure-Learn framing for testing the population-count assumption with cheap experiments before committing to the SOM as the foundation of the business plan. Reference at theleanstartup.com.
- CB Insights. State of the Markets report series. Industry-specific market sizing for technology categories, published with named methodology; a defensible secondary source for category-level TAM backdrop. Reference at cbinsights.com.
- Pitchbook and Crunchbase. Comparable-Set Revenue Data. The public revenue disclosures and funding rounds the article uses as comparable-set inputs for the per-customer revenue figure. Reference at pitchbook.com and crunchbase.com.
The source audit deliberately excludes the popular "$X trillion market by 2030" analyst forecasts without named methodology, copied headline TAMs from third-party pitch decks with no traceable source, and invented per-customer math that multiplies a category size by a percentage with no basis. The article cites the US Census Bureau and the BLS as primary sources for US-based sizing because both are free, both publish the methodology, and both can be reproduced by any reader. For Yibud's broader evidence policy, see Sources & references.
Continue learning
Where to go from here
These pieces are grouped by topic, not publication date — pick the one that matches the question you are working on right now.
More in ValidationSee all topics →
Validation Guide
How to Write a Smoke-Test Landing Page in 60 Minutes
The shortest path from a hypothesis to a real conversion number — page structure, headline formula, the four signal lines, the seven things to leave off, the 60-minute build checklist, and how to read the conversion data at small N.
14 min read
Validation Guide
Fake Door Test: How to Validate Demand Without Writing a Single Line of Code
A complete playbook for the fake door (painted door) test — the cheapest credible demand test in the startup validation toolkit. Includes the original methodology from Alistair Croll & Benjamin Yoskovitz's Lean Analytics, the six-step process, the Unbounce 2024 benchmark that defines what counts as a 'good' smoke-test conversion rate, the two documented case studies (Buffer, Dropbox), the seven failure modes that produce false positives, and what the test cannot prove.
16 min read
Validation Guide
How to Validate Demand Before Coding (Without Building the Wrong Thing)
A step-by-step playbook for non-technical and solo founders who need to know whether anyone will pay for an idea before writing a line of code — including the five pre-code evidence checks, three labeled hypothetical examples, the seven-question rule that decides 'build now vs. validate more,' and the failure modes that lead founders to mistake enthusiasm for demand.
14 min read
Monetization Validation
How to Price a New SaaS Product: Three Decisions, Four Questions, and the Only Pricing Test That Survives Launch
The three pricing decisions a new SaaS founder actually owns (model, value metric, price points), the four Van Westendorp questions that produce real willingness-to-pay data, and the eight-week playbook that turns the answers into a defensible price.
18 min read
Validation Guide
Startup Failure Analysis: How to Read a Failed Startup the Right Way
Startup failure analysis is the discipline of converting documented startup failures into testable assumptions. Five real cases — Quibi, Webvan, Juicero, Homejoy, WeWork — analyzed with a seven-rung framework, then turned into the validation experiments a solo founder can run this week.
26 min read
Validation Guide
Lean Startup Validation: How To Choose, Design and Interpret Validation Experiments
How to use Eric Ries's Build-Measure-Learn loop to pick the next validation experiment — including the five-question decision framework, decision-threshold thinking, three labeled hypothetical examples, and the failure modes that distort the loop before the founder notices them.
17 min read
Validation Guide
Problem-Solution Fit: How to Know Your Solution Solves a Real Problem Before You Build It
Problem-solution fit is the state where a defined customer has a defined problem and a defined solution would meaningfully improve their situation. The five rungs of evidence that prove the fit is real, the six failure modes that show it isn't, and the 14-day validation framework that separates real pain from polite enthusiasm.
22 min read
Validation Guide
MVP Validation: How to Test a Minimum Viable Product Before You Build It
MVP validation is the discipline of testing the smallest useful version of your idea before committing months to a full build. Six methods, a five-step framework, and the seven failure modes that show up before the founder notices them.
20 min read
Validation Guide
Product-Market Fit Validation: How to Know You Have It Before You Announce It
Product-market fit is observed, not declared. The four signals that show you have it, the seven failure modes that show you do not, and a 30-day validation framework that separates real retention from paid acquisition.
22 min read
Monetization Validation
Willingness to Pay Validation: The Framework, the Signal Ladder, and Why Interest Is Not Payment
What willingness to pay actually means, the five-rung signal ladder from polite words to real money, and the Problem → Customer → Value → Price → Payment framework every pricing experiment is a sub-test of — before you write code.
14 min read
Validation Guide
How to Validate an API Startup Idea Before You Build It
The API- and developer-tool-specific tests for technical buyers, integration cost, trust, documentation prototypes, and design-partner pilots — before you write the first endpoint. A practical handbook for API founders, SDK builders, infrastructure product teams, and developer-tool indie hackers.
18 min read
Validation Guide
How to Validate a B2B Startup Idea Before You Build It
The B2B-specific tests for buying committees, procurement, ROI proof, founder-led sales, and the manual pilot — before you write code. A practical handbook for SaaS founders selling to businesses, enterprise software teams, and technical founders.
17 min read
Validation Guide
How to Validate a Chrome Extension Idea Before You Build It
The browser-extension-specific tests for Manifest V3 fit, Chrome Web Store policy, distribution outside store search, willingness to pay, and unlisted pre-launch testing — before you ship a packaged extension. A practical handbook for indie hackers, SaaS founders, AI tool builders, and browser extension developers.
16 min read
Validation Guide
How to Validate a Mobile App Idea Before You Build It
The mobile-app-specific tests for problem, retention, onboarding, distribution, and willingness to pay — before you ship a binary to the App Store. A practical handbook for consumer, productivity, lifestyle, health, education, and local-service apps.
17 min read
Validation Guide
How to Validate a Marketplace Startup Before You Build It
The marketplace-specific tests for supply, demand, liquidity, take rate, and two-sided interviews — before you build the platform. A practical handbook for B2B, consumer, local, creator, and talent marketplaces.
17 min read
Validation Guide
AI Startup vs SaaS Startup: How Validation Is Different
Why AI startups need workflow, output-quality, and dependency tests on top of every SaaS validation question — and the cheapest experiment that proves each one before you build.
18 min read
Validation Guide
How to Validate an AI Startup Idea
Validate an AI startup idea by testing workflow demand, output quality, pricing, model dependency, distribution, and defensibility before building.
21 min read
Validation Guide
How to Validate a SaaS Idea Before You Build It
The five recurring-revenue assumptions that decide whether a SaaS product survives month six, and the cheapest experiment that tests each one — before you write code.
16 min read
Monetization Validation
How to Test Willingness to Pay Before You Build: Six Experiments, Four Cases, and a Seven-Day Plan
Six pricing experiments you can run this week, four documented startup cases that show the pattern, and a seven-day Stripe-checkout plan that asks for money — before you build.
14 min read
Validation Guide
Startup Validation Checklist: Before You Build
21 concrete checks across four validation stages — problem, customer, business, execution — with how to test each one, the common mistake to avoid, and a printable summary.
17 min read
Validation Guide
How to Validate a Startup Idea Before You Build
The four assumptions every startup depends on, the four questions that test them, and the cheapest experiments that produce evidence in 2–4 weeks — before you build.
17 min read
See every article on startup validation in one place.
Open the Startup Validation hub →