Manual-behind-the-curtain MVP
Wizard of Oz MVP — Fake the Automation, Deliver by Hand
A Wizard of Oz MVP is not a half-built app. The customer uses something that looks like a product — a form, a generate button, an email that arrives on time — while you produce the output by hand. The question is whether that experience is worth automating, not whether you can already ship the factory.
Last updated: September 25, 2026
Direct answer
What is a Wizard of Oz MVP?
Eric Ries, in The Lean Startup (Crown Business, 2011), treats the Wizard of Oz test as a learning vehicle: the customer believes they are using the product; the founder operates the machinery out of sight. That is the opposite of a concierge MVP, where the founder is openly the service. The name is older than lean startup. J.F. Kelley described the HCI “Wizard of Oz” method in ACM Transactions on Office Information Systems (1984): a human simulates a computer system so researchers can study the experience before the system exists. Ries applied the same curtain to startups.
Key takeaways
What to remember
- The customer thinks the product is real. You are the backend. If they know it is you, you are running concierge, not Wizard of Oz.
- Use it when the expensive unknown is the experience — format, latency you can actually keep, whether they come back — not “will anyone click a landing page.”
- Write pass, fail, and the stop-faking rule before the first customer arrives. A bar you invent after a messy week is not a result.
- Deliver the promised outcome. Hiding labor is a test. Advertising a capability you cannot produce even by hand is not.
- Stop faking when the same path repeats without improvisation — or when you cannot keep the SLA without hiring. Faking forever is an unpaid ops job.
Which test
When to use Wizard of Oz vs concierge, smoke test, or fake door
Founders (and a lot of blog posts) use these names as synonyms. They are not. Pick the cheapest test that answers the question you actually have. The four sit in different places.
Wizard of Oz — the curtain is up
You can produce the outcome by hand at a latency the customer would believe from software. You need to know whether they will use, repeat, or pay for that experience — not whether they will sit with you on a call. The front looks like a product. The back is you, a spreadsheet, and a timer.
Concierge — they know it is you
You need to watch the workflow with the customer in the room. Hiding the labor would be dishonest (they are buying “a person who does this”), or you cannot keep a believable automated SLA. Concierge is openly manual on purpose.
Read the concierge playbook →Smoke test — no delivery yet
The expensive unknown is still demand. You do not know if strangers will reach for the concept. A landing page plus one CTA is enough. Do not stand up a fake backend to count clicks.
Read the smoke-test playbook →Fake door — one button, existing product
People already use something you shipped. You want to know if they will reach for one addition. Put a dead button where they already work. Do not rebuild the company as a Wizard of Oz to test a menu item.
Read the fake-door playbook →This week
Seven steps to run a Wizard of Oz test this week
Keep the front thin and the backstage script boring. A clever UI that you cannot operate on Tuesday night is not a test.
- 1
Write one falsifiable bet
Finish this sentence before you build a form: “People who [role] and recently [pain] will [complete the job / request a second run / pay X] for [named outcome] within [timebox] when the front looks automated.” If you cannot finish it, you are not ready.
- 2
Freeze pass, fail, and the stop-faking rule
Who counts as a qualified customer. What action you will count. How many you will serve before you read the result. The calendar date you will decide. And the line that ends the curtain: same path, no improvisation — or you miss the SLA twice. Write it on the same page as the bet.
- 3
Design the thinnest honest front
One input, one output, one turnaround you can keep while you are awake. A Typeform, a “generate” email, a Notion page that looks like an app. Do not advertise 24/7, instant, or “no human reads this” if any of those is false.
- 4
Write the backstage script
Tools, minutes, and what you will not do. Time the first dry run on yourself. If the happy path needs a custom judgment you cannot describe in five steps, you do not have a product candidate yet — you have a service.
- 5
Recruit a handful of right customers
Three to five ICP-matched people who already feel the pain. Prefer strangers who would eventually pay. Friends inflate “it felt like magic.” This is not a launch. Do not buy a burst of traffic you cannot fulfill.
- 6
Deliver on the SLA and log every exception
Hit the turnaround you advertised. After each job, write: minutes spent, what you improvised, whether they used the output, whether they asked for another. The improvisation log is the automation backlog.
- 7
Follow up, then decide against the bar you wrote
Ask for a second cycle and a price condition. Ask what they thought was automatic. Compare the week to step 2 — not to a vibe. Keep the curtain, disclose and switch to concierge, automate one repeated step, or stop.
Pass / fail
Set the thresholds before the first customer
A Wizard of Oz week without a pre-written bar becomes a story you tell yourself about how busy you were. Write the five lines below on the same day you write the front. This page will not invent a universal conversion rate. Channel, offer, and how hard the job is are not interchangeable.
- Qualified customer: who counts (role + recent pain). Everyone else is discarded, not averaged in.
- Conversion event: the one action you will count — used the output, requested a second run, or paid. A compliment is not an event.
- Sample: how many qualified people you will serve before you read the rate. A handful you can fulfill beats a waitlist you will ghost.
- Pass line and fail line: two outcomes you can live with. “Interesting” is not a line.
- Timebox and SLA: the calendar date you will decide, and the turnaround you must hit for the test to stay valid.
Signal still ranks the same way: a second paid cycle beats a thank-you; a used artifact beats a click. Compare yourself to the bar you wrote — not to someone else’s dashboard screenshot.
When to drop the curtain
When to stop faking
The curtain is temporary on purpose. Write the stop rule before you start so “one more customer” does not become your job.
Stop faking and automate when the path repeats
You can describe the happy path without improvising. The same painful step shows up for more than one customer. Automate that step first — not the rare flourish that made one person smile.
Stop faking and disclose when the SLA is the product
You missed the advertised turnaround twice, or you would have to stay awake to keep the lie. Either narrow the promise or tell them it is manual (that is concierge) before you take the next job.
Stop faking when they are buying you, not the product
They only value “a smart person on the other end.” That is consulting. Keep it as concierge if you want the learning, or walk away. Do not automate a personality.
Stop the experiment when nobody comes back
Qualified people take the first job and do not request a second, do not use the output, and will not name a price. That is evidence against this experience — not a reason to build a faster backend.
Worked example
One named history, one week you can copy
The Zappos story is a founding case, not a conversion benchmark. The week below is illustrative — names and prices are invented so you can see the setup, not a result you should copy as a number.
Zappos, 1999 — the storefront was real; the warehouse was not
Nick Swinmurn photographed shoes in local stores and listed them online. When an order arrived, he bought the pair at retail and shipped it. The customer thought they were buying from an online shoe store. Eric Ries recounts this in The Lean Startup (2011) as validated learning: test whether people will complete the job before you build the factory. No inventory system, no claimed fill-rate. The lesson is the curtain, not a percentage.
This week — an “AI research brief” you write by hand
Illustrative only. A founder wants to sell a $49 brief that turns a URL plus one question into a two-page memo for independent consultants. The expensive unknown is not “will anyone click.” It is “will they use a memo that arrives in four hours and ask for another.”
- Monday: write the bet and the fail line. Front = a form (URL + question + email). SLA = four hours during weekday 9–6. Backstage = the founder, a browser, and a timer. Pass = three of five qualified consultants use the memo and request a second. Fail = fewer than two use it.
- Tuesday: dry-run one brief on your own last client note. If it takes three hours, the SLA is a lie — cut the memo or the hours before anyone arrives.
- Wednesday–Friday: five ICP-matched consultants, not a Product Hunt launch. Log minutes, improvisations, and whether they forwarded or filed the memo.
- Sunday: decide against Monday’s bar. If they used it and asked for another, automate the outline — not a fake “AI engine” page. If they only liked talking to you on Slack, you ran concierge by accident.
Checklist
Copy-ready run checklist
- Bet written in one sentence with a timebox
- Pass line, fail line, and stop-faking rule written before recruitment
- Front has one input, one output, one SLA you can keep awake
- No claim you cannot produce by hand (24/7, instant, “no human reads this”)
- Backstage script timed on a dry run
- Customer count capped at what you can fulfill this week
- Log of minutes, exceptions, and whether they used the output
- Follow-up for a second cycle and a price — then a one-sentence decision
Common mistakes
What usually ruins the test
Building a fake UI that takes longer than doing the work
You spend four days on a dashboard and one hour on the job. The test is the job. A form is enough.
Advertising latency or privacy you cannot keep
“Instant” and “no human sees your data” are product claims. If you are the human, those claims are false. Missed SLAs also invalidate the experience you were trying to measure.
Charging for a finished product you cannot refund or deliver
A small fee for a delivered brief is a test. Taking a card for “the platform” you do not have is a different legal and ethical posture. If you cannot deliver or refund this week, do not take the card.
Never writing the pass/fail line
A busy week feels like traction. It is not. If you did not write the bar on Monday, you do not have a result on Sunday.
Calling it Wizard of Oz when they know it is you
If you are on the call, in the inbox as yourself, or “white-glove onboarding” as the founder, that is concierge. Use the concierge playbook. Mixing the two names hides whether the experience survived without you.
Faking forever
No timebox, no automation candidate, no disclosure plan. That is an unpaid service business with a product landing page.
Honesty
The curtain is a test, not a business model
Wizard of Oz is more ethically loaded than concierge because the customer does not see the labor. The line is simple: deliver the outcome you promised; do not sell a capability you cannot produce even manually; plan how the curtain ends.
Promise an outcome, not a factory
“You get a two-page brief by 4pm” is a promise you can keep by hand. “Our AI never sleeps” is a claim about a system you do not have.
Do not take money you cannot refund or fulfill
If you charge, deliver this week or refund. A deposit on a delivered job is a stronger signal than a waitlist. A subscription to software that does not exist is not this test.
Have a disclosure or graduation plan
When they ask “is this automated?”, answer. When you automate, the experience should not get worse. Customers who felt cheated when the “app” later needed them to wait longer were not validated — they were surprised.
Collect the minimum
Take the file or URL you need to do the job. Do not harvest credentials, inboxes, or “training data” you will not use.
Continue or stop
After the timebox
Keep going when
Qualified people use the output, request another cycle, and tolerate rough edges that are not the SLA. Next: automate the highest-frequency step, or disclose and run concierge if you still need to watch the workflow.
Pause or pivot when
Nobody wants a second cycle, they only valued talking to you, or every job is a unique custom path. Do not “build the real engine anyway.” Rewrite the outcome or the ICP.
Sources
Where these ideas come from
- Eric Ries, The Lean Startup (Crown Business, 2011) — Wizard of Oz and concierge as learning vehicles before automation; the Zappos storefront-before-warehouse story as validated learning, not as a conversion rate.
- J.F. Kelley, ACM TOIS (1984) — The HCI “Wizard of Oz” method: a human simulates the system so you can study the experience before the system exists. The name predates lean startup.
- Alberto Savoia, Pretotype It (2011) — The Mechanical Turk / Wizard of Oz pretotype: a real-looking front, manual work in the back, built to test whether the experience is the right it.
With Yibud
Use the report to decide whether the curtain is the right next move
Yibud is a free, no-signup startup idea validator. Scores come from a deterministic rule engine — not from a language model. Optional AI only polishes prose. If the weakest dimension is delivery or “will they use this experience?”, a Wizard of Oz week is often cheaper than a build. If the weakest dimension is still demand, run a smoke test first.
Analyze my idea →FAQ
Questions founders actually ask
Is a Wizard of Oz MVP the same as a concierge MVP?
No. In a concierge MVP the customer knows a human is delivering the service. In a Wizard of Oz MVP the customer believes the product is automated. Same labor, different question: concierge watches the workflow with them; Wizard of Oz tests whether the experience holds when they think a system did it. Use the concierge guide when honesty about the hands is the point.
Is this just a smoke test with extra steps?
No. A smoke test is a whole-concept landing page plus one CTA — no delivery. Wizard of Oz delivers the outcome. If you have not yet seen anyone reach for the concept, you do not need a fake backend. Use the smoke-test guide first.
Is this the same as a fake-door test?
No. A fake door is a feature-level button or menu inside a product people already use. The conversion event is a click, then you explain the feature is not ready. Wizard of Oz is a whole experience you actually fulfill by hand. Use the fake-door guide for in-product reach.
Is it unethical to hide the labor?
It is unethical if you claim a capability you cannot produce, take money you cannot refund, or vanish. It is a normal early-stage experiment if you deliver the promised outcome, keep the SLA, collect the minimum, and have a plan to disclose or graduate. Do not advertise “no human in the loop” when you are the loop.
How many customers do I need?
Start with a handful you can serve at the advertised turnaround. Depth beats a launch you cannot fulfill. There is no universal “statistically valid” number for this stage — write your own sample on Monday and keep it.
Do I have to charge?
Not on day one, but a price conversation should happen in the same week. Payment — even small — separates politeness from priority. Only charge if you can deliver or refund.
When do I start coding?
When the same painful step repeats across customers and you can describe the happy path without improvising every time. Automate that step. Do not start with a general “AI engine.”
How does this connect to Yibud?
Run the analyzer if you do not yet know which assumption is most expensive. The report will not operate the curtain for you. It will tell you whether you should be counting demand, talking to people, or delivering an experience by hand this week.
Know which risk to test before you automate
Yibud scores the weak dimensions with a deterministic engine, then you pick the experiment — a conversation, a smoke test, or a week behind the curtain.
Generate a Yibud report →