A Wizard of Oz MVP hides a human doing the work behind a software-looking front end, so users believe they're using an automated product. A Concierge MVP does the same manual work in the open, with users knowingly paying a human. Both let you test whether people want the outcome before you build the automation that produces it.

Quick Answer: Use Wizard of Oz when you need honest reactions to a system that feels automated (pricing, trust, adoption behavior). Use Concierge when you need deep qualitative insight and can be transparent that a human is delivering the service by hand. Pick based on whether the illusion of automation is itself part of what you're testing.

Both are forms of what Eric Ries, in The Lean Startup, called building the smallest thing that lets you run a validated-learning loop — except here the "thing" has no automated engine at all. You are the engine. That distinction matters more than most teams treat it: it determines whether you learn about the value proposition or about your interface design, and conflating the two wastes the exact runway an early-stage MVP is supposed to protect.

What Separates Wizard of Oz from Concierge MVPs

The core difference is disclosure, not effort. In a Wizard of Oz MVP, users interact with what looks like a functioning product — a form, a dashboard, a matching algorithm — while a human executes every step behind the scenes, unseen. In a Concierge MVP, there's no curtain: users know a person is doing the work, often on a call or over email, and that's part of the deal.

The classic reference point is Zappos. Nick Swinmurn tested whether people would buy shoes online, sight unseen, by photographing shoes in local stores, posting them, and personally walking to the store to buy and ship any pair that sold — no inventory, no warehouse, no automated fulfillment. That's Concierge: manual, but not disguised as anything else.

Drew Houston's early Dropbox test worked differently. Before building the syncing engine, he made a screen-recorded demo video that walked through folder sync as if it already worked, to gauge signup demand for a product that, at that moment, barely existed behind the curtain. It leaned Wizard-of-Oz in spirit — presenting the finished experience before the engine existed — even though it used a video rather than a live disguised backend.

The Decision Table

DimensionWizard of Oz MVPConcierge MVP
User awarenessBelieves it's automatedKnows a human is helping
Best for testingTrust in the system, interface flow, willingness to act on machine outputDepth of the problem, service quality, willingness to pay for the outcome
Typical scaleA handful to a few dozen users, one at a timeVery small, often single-digit users, high-touch
Risk if discoveredHigh — breaks trust, feels deceptiveNone — disclosure is the design
Learning speed per userSlower (must maintain illusion)Faster (can ask directly, iterate live)
Ethical barRequires a debrief/disclosure planBuilt-in from the start

The practical takeaway: if the thing you're most unsure about is whether users will trust and act on an automated recommendation — a matching algorithm, a personalized feed, an AI-generated plan — you need the illusion of automation intact, so Wizard of Oz is the right shape. If you're unsure whether the underlying job is even worth solving for someone, Concierge gets you there faster and more honestly.

When Wizard of Oz Fits

Wizard of Oz MVPs fit when the automation itself is the risky part of your product hypothesis — when you need to know if users will act on, trust, or pay for output that looks machine-generated, before investing in the machine. Reveal too early or too crudely, and you contaminate the very trust signal you're trying to measure.

A well-known example: many early recommendation and matchmaking products — including versions of Aardvark, the Q&A-routing tool later acquired by Google — used humans to route questions to the right expert respondent behind a chat interface that felt automated, before any of that routing logic was actually coded. Users experienced a system; the "algorithm" was a person reading and dispatching by hand.

What Makes a Good Wizard of Oz Candidate

  1. The output is evaluable without knowing how it was produced — a recommendation, a match, a generated plan, a scheduled slot.
  2. A human can plausibly produce that output within the latency users will tolerate — minutes, not days, or the illusion collapses.
  3. You have a real hypothesis about automation-worthiness, not just curiosity — write it down using the same discipline as any testable product hypothesis: "If we show users a matched vendor within 10 minutes, at least 30% will book a call."
  4. You can operationally scale the manual step to your test cohort size without collapsing under volume — Wizard of Oz breaks down fast past a few dozen concurrent users.

The Manual-Matching Example

Say you're building a B2B vendor-matching tool: a buyer describes a need, the "system" instantly suggests three vendors. Behind the scenes, a human — you, a cofounder, an ops hire — reads the request, searches a spreadsheet of vendors, and manually populates the three results within an agreed SLA (say, 15 minutes).

You're not testing the matching algorithm. You're testing:

  • Whether buyers trust a same-session match enough to click through
  • Whether the three-vendor framing is the right choice architecture
  • Whether the speed of response, not the quality of the algorithm, is what drives conversion
  • Whether buyers escalate to a real conversation, or bounce after seeing suggestions

If buyers don't click through even when a sharp human hand-picks the best possible matches, no amount of matching-algorithm sophistication will fix that. That's the entire point of faking the backend — it isolates the value proposition from the automation risk.

When Concierge Fits Better

Concierge MVPs fit better when you need rich qualitative signal about a problem you don't yet fully understand, and disguising the human would only get in the way of asking direct questions. It trades scale for depth — you serve very few users, but you learn far more per user than any illusion allows.

Because disclosure is built in, you can ask the user directly: "What almost stopped you from using this? What would you have paid? What did you expect next?" That's a live interview embedded in the delivery of the actual service — closer to running structured discovery than running an experiment. This maps well onto jobs-to-be-done interviewing: you're not asking people to imagine a job, you're watching them hire you for it in real time, manually, and narrating your own reasoning back to them.

A Concierge Walkthrough

Consider a meal-planning-for-diabetics idea. Before building any nutrition-matching engine, the founder personally messages five users each week: asks about symptoms, blood sugar logs, and preferences, then manually assembles a meal plan and sends it as a PDF, following up by text.

That manual loop surfaces things a Wizard-of-Oz test would hide by design: users edit the plan out loud ("I'd never eat this for breakfast"), reveal decision criteria never captured in a signup form assess whether five days is often too long a planning horizon, and expose which parts of the "product" people actually value — the meal ideas or the accountability check-in. Concierge MVPs are better at answering "what should we even build" than "will people trust what we build."

Deciding When to Graduate from Manual to Built

Graduate from manual to automated when three conditions hold together: demand is validated and repeatable, the manual process has stabilized into a describable, repeatable procedure, and the unit economics of staying manual are actively degrading as volume grows. Automating before all three hold usually means automating the wrong thing.

The Graduation Checklist

SignalStill manualReady to build
Demand patternSporadic, unpredictable usersRecurring requests, a describable segment
Process shapeYou're still improvising each timeYou could write an SOP a new hire could follow
Time cost trendRoughly flat as you learnRising per-user as volume grows
What breaks firstNothing yet — you can absorb loadYou, your ops person, or your inbox
Confidence in what to automateStill guessing at the algorithmYou could specify the logic in plain rules

Two of the same forces power/simplicity ideas from Ash Maurya's Running Lean and the classic "concierge before code" advice attributed to the Lean Startup and customer-development lineage (Steve Blank, Eric Ries): don't build the engine until you can specify, in advance, roughly what decisions the engine needs to make. If you still can't articulate the matching logic after 40 manual matches, more manual reps — not more code — is the right next investment.

Before committing engineering time, run the numbers the same way you would for any A/B test setup: what's the minimum sample of manual reps needed before a pattern is real rather than noise, and do you have enough signal to trust it? A related discipline — knowing the actual statistical-significance numbers a PM needs — applies just as much to "does this manual process show a stable conversion rate" as it does to a formal split test.

A Graduation Sequence That Works in Practice

  1. Run manual with 5-20 users, tracking conversion, satisfaction, and your own time cost per case.
  2. Write the SOP — the exact steps you take, in order, as if handing it to a new hire tomorrow.
  3. Partially automate the bottleneck only — usually the slowest or most repetitive single step, not the whole flow.
  4. Keep a human in the loop for edge cases while the automated core handles the common path.
  5. Fully automate once edge cases are rare enough that the human-in-the-loop step costs more than it protects against.

This sequence — full manual, then partial automation with a human safety net, then full automation — mirrors how many real "AI-powered" products actually launched: a human reviewing or overriding model output long after the marketing already said "automated." That's a legitimate, common architecture, not a dirty secret, as long as it's not misrepresented to users as fully autonomous when it materially isn't.

Where This Fits in Your Broader Experimentation Practice

Wizard of Oz and Concierge MVPs are experiment designs, not one-off tricks — they belong in the same disciplined loop as any other test in your experimentation practice. Each one needs a hypothesis, a defined sample, a decision rule for what "worked" means, and an honest post-mortem — otherwise you'll rationalize a mediocre result into a green light.

Write the hypothesis before you run a single manual match. "If we manually deliver X within Y minutes, at least Z% of users will do [target action]" is the same shape of statement you'd write for any testable product hypothesis — the fact that a human is standing in for software changes nothing about the rigor the hypothesis itself needs.

Mapping where the manual illusion has to feel seamless is where a customer journey view earns its keep. Prodinja's Customer Journey tool is built around an emotion curve across the steps a user takes — and for a Wizard of Oz test specifically, that curve is a useful lens for identifying exactly which moment (the wait for a "match," the framing of a "recommendation") has to feel effortless for the illusion, and therefore the test of the real user promise, to hold.

Key Takeaways

  • Disclosure is the real dividing line between Wizard of Oz (user believes it's automated) and Concierge (user knows a human is helping) — not how much manual effort is involved.
  • Wizard of Oz tests trust in automation itself — use it when the risky assumption is whether users will act on machine-feeling output.
  • Concierge MVPs generate richer qualitative signal because you can ask users directly what they need, in real time, without maintaining an illusion.
  • Zappos and early Dropbox are the reference cases — one manual-fulfillment Concierge test, one demo-video Wizard-of-Oz-style test of latent demand.
  • Graduate to automation only when demand is repeatable, the process is describable as an SOP, and manual costs are rising — not on a fixed calendar date.
  • A human-in-the-loop stage between manual and fully automated is a legitimate, common step — not a compromise to hide from users, provided it's represented honestly.
  • Treat both approaches as real experiments: write the hypothesis first, define the sample and decision rule, and be honest in the post-mortem about what the result actually shows.

Frequently Asked Questions

Is a Wizard of Oz MVP ethical if users don't know a human is involved?

Yes, when scoped correctly — you're testing the product experience the same way you'd test any prototype, and users aren't being financially harmed or deceived about anything material like pricing or data use. Best practice is a plan to disclose the mechanism afterward if participants are identifiable, and to never fake something with real safety, legal, or financial consequences.

How many users do you need for a Wizard of Oz or Concierge test?

Far fewer than a statistical A/B test — often 5 to 30 users is enough to see a clear behavioral pattern, since the goal is directional signal on a value proposition, not p-values. If you truly need statistical confidence at scale, you've likely outgrown the manual-test stage and should move toward a proper A/B test.

What's the difference between a Concierge MVP and just doing customer interviews?

A Concierge MVP delivers the actual outcome by hand, so you observe real behavior and willingness to pay, not just stated intent. Interviews capture what people say they'd do; Concierge captures what they actually do when the service is real, even if the delivery mechanism behind it is a person instead of software.

Can a Wizard of Oz test work for a hardware or physical product?

Yes, though it's less common — teams have used manually operated back-ends behind seemingly automated physical interfaces (a "smart" kiosk actually staffed by a hidden operator, for instance) to test interaction design before investing in embedded engineering. The same rule applies: disguise only the mechanism, never material facts like safety or pricing.

When should you skip Wizard of Oz or Concierge entirely and just build?

Skip manual testing when the automation itself is trivial to build and the real risk is something else entirely, like distribution or regulatory approval — faking a backend only pays off when the engine is the expensive, uncertain part. If building the real thing is cheaper than convincingly faking it, build the real thing.