Test the assumption that would kill the project if it's false — not the one your team already feels sure about. Confidence tracks how familiar an idea feels, not how true it is, which is why teams test the safe assumption while the one that could sink the bet sits untested.
Quick answer: Score every assumption on two axes — impact if it's wrong, and how much real evidence you already have. The one with high impact and low evidence is your riskiest assumption. Test that one first, this week, with the smallest experiment that could prove you wrong.
Why Confidence Keeps Teams Testing the Wrong Assumption First
Teams default to testing whatever assumption is easiest to check, not the one that matters most, because confidence is generated by familiarity and repetition rather than accuracy. The assumption you've discussed in ten meetings feels solid simply because you've discussed it ten times — that feeling has nothing to do with whether it's true.
Daniel Kahneman's research on the planning fallacy and overconfidence (with Amos Tversky) found that people systematically rate their own plans as more likely to succeed than the evidence supports, especially once they've spent a long time inside the plan. The more a team has invested in an idea, the more certain it feels — regardless of what's actually been tested.
This produces what's sometimes called the streetlight effect — searching where the light is good, not where the keys actually fell. Two assumptions rarely carry equal testing cost:
- Cheap to check: whether a color scheme lands well, or whether a headline tests better than another — a five-question survey, done in an afternoon.
- Expensive to check: whether a health system will integrate with a new API, or whether a buyer will sign at the price modeled in the deck — real effort, real risk of an awkward "no."
Convenient wins by default, every time, unless a team deliberately overrides it. This bias has a real cost. CB Insights' recurring analysis of startup postmortems consistently finds "no market need" among the single most-cited reasons founders give for failure — roughly a third of cases, year after year. Most of those teams didn't skip testing entirely. They tested the wrong thing, confidently, for months.
The question isn't "what can we check easily?" It's "what, if wrong, breaks the whole plan — and how sure are we, really?"
This is one specific move inside a wider practice, covered in full in a complete guide to adversarial thinking in product work: deliberately arguing against your own plan before the market does it for you. Assumption prioritization is the narrowest, most tactical piece of that discipline — and often the one teams skip first.
Assumption Mapping: Plot Impact-If-Wrong Against Current Evidence
An assumption map is a simple 2x2: one axis scores how bad it would be if an assumption is wrong (impact), the other scores how much real evidence already supports it (evidence). Plot every assumption behind your bet, and the one in the high-impact, low-evidence corner is the assumption to test first.
This isn't a new idea. David J. Bland and Alexander Osterwalder, in Testing Business Ideas — the Strategyzer follow-up to Business Model Generation — formalize this as an Assumptions Map, scoring each assumption on importance and evidence so teams stop treating every belief as equally worth testing. Assumption mapping, done this way, is a product discipline any team can run before committing a roadmap slot, not a one-off workshop exercise.
Once mapped, four distinct quadrants emerge, and each demands a different response:
- Leap-of-faith (high impact, low evidence): Untested and lethal if wrong. This is your riskiest assumption — test it now.
- Validated bet (high impact, high evidence): Already load-bearing and already checked. Keep monitoring; don't waste a cycle retesting what you've earned confidence in.
- Curiosity (low impact, low evidence): Interesting, unproven, but survivable either way. Park it and move on.
- Safe assumption (low impact, high evidence): Settled and inconsequential. Stop spending meeting time defending it.
Laid out as a grid, the priority becomes obvious at a glance:
| Low Evidence | High Evidence | |
|---|---|---|
| High impact if wrong | Leap-of-faith — test first | Validated bet — proceed, monitor |
| Low impact if wrong | Curiosity — park it | Safe assumption — stop testing |
Only one quadrant deserves your next test cycle. Everything else can wait, get lightly monitored, or simply get logged and left alone — the map's real job is to stop every assumption from competing for the same attention.
Evidence quality matters as much as evidence quantity. Rob Fitzpatrick's The Mom Test documents how customer conversations routinely produce false confidence — leading questions, polite lies, and hypothetical enthusiasm all get logged as "evidence" when they're closer to noise. An assumption map only works if the evidence column reflects testing that could have proven you wrong, not testing that was structurally incapable of doing so.
Two of the more honest sources of evidence for a customer-facing assumption are a customer journey map built from what people actually do, and a clear Jobs to Be Done analysis of what they're actually trying to accomplish. Both are harder to fake than a friendly conversation, because both are built to surface contradictions, not agreement.
The Riskiest Assumption Test (RAT) Loop, Step by Step
A Riskiest Assumption Test, or RAT, is a small, fast experiment designed to falsify the single assumption sitting in your map's kill zone — not to validate the whole product idea. The loop runs in days: name the assumption, design the smallest test that could prove it false, run it, then decide to kill, pivot, or proceed.
The term comes from Tristan Kromer, a Lean Startup practitioner who argued that MVP had become a confusing, overloaded label, and proposed RAT as a sharper substitute: build the smallest thing that tests your riskiest assumption, not the smallest version of the whole product. It extends Eric Ries's original idea of leap-of-faith assumptions in The Lean Startup — the handful of beliefs a business plan can't survive being wrong about.
The loop itself is short enough to run every time a team is about to commit real budget:
- List every assumption the bet depends on. Technical, market, pricing, behavioral, operational — don't pre-filter yet.
- Score each on impact and evidence. Use the 2x2 above, and be honest about what actually counts as evidence.
- Name the single riskiest assumption. Highest impact, lowest evidence wins — resist the pull toward whichever is easiest to test instead.
- Design the smallest test that could prove it false. Not confirm it — falsify it. If every plausible outcome confirms your belief, the test is worthless.
- Set a kill threshold before you run it. Decide in advance what result means "stop" — deciding after seeing the data is how confidence creeps back in.
- Run it in days, not sprints. A
RATthat takes six weeks has usually turned into a small build in disguise. - Kill, pivot, or proceed — then re-map. The loop repeats; it doesn't end after one pass.
Step 1 is where most teams under-list, because they're too close to their own plan to see what they're assuming. Adversarial techniques exist specifically to surface those blind spots: a four critics premortem panel forces multiple failure vantage points before a launch, a skeptical engineer critique is built to expose the feasibility assumptions hiding inside "we can build this," and a revenue hawk critique does the same for "someone will pay for this." Run one of these before step 1, and the assumption list gets longer, and considerably more honest.
Marty Cagan's framing of product risk — value, usability, feasibility, and business viability, laid out across INSPIRED and EMPOWERED — maps cleanly onto this loop too. Most teams instinctively over-test usability (easy, cheap, comfortable) and under-test viability and value (uncomfortable, slow, expensive) — the exact convenience bias this whole exercise exists to correct.
A One-Week Walkthrough: How a RAT Loop Killed a Bad Bet
Picture a common pattern in enterprise software: a team believes finance buyers will pay a premium for an automated compliance report and builds a quarter's roadmap on that belief. A one-week RAT loop, run before a single sprint is staffed, would surface the flaw for a fraction of the cost of building it first.
Here's how the week could play out:
- Day 1 — List and score. The team maps three assumptions: technical (can the report be generated accurately), behavioral (will finance actually open it), and pricing (will they pay extra for it). Pricing scores highest impact, lowest evidence — nobody outside the sales team had asked a real buyer.
- Day 2 — Design the smallest falsifiable test. Instead of building the feature, the team mocks up a single-page version of the report with a real price attached, and sets the kill threshold in advance: fewer than three of ten target buyers willing to discuss budget kills the assumption.
- Days 3–4 — Run it. Sales walks the mockup past ten warm accounts, using open, non-leading questions in the spirit of The Mom Test rather than "would you like this?"
- Day 5 — Decide. One buyer engages seriously; the rest already get comparable data from an auditor, or don't rank it as a budget priority. The pricing assumption fails its own threshold.
The team kills the premium-report angle before a single engineer is staffed on it, and redirects the quarter toward the piece the mockup actually validated: finance wanted the report inside their existing workflow, unpriced, as a retention feature rather than a paid upsell.
Nothing here required six weeks or a shipped feature. It required naming the assumption honestly, setting a kill threshold before ego got involved, and treating "that's a good question, let me get back to you" as the answer it actually was.
Where Assumption Testing Fits Your Broader Adversarial-Thinking Toolkit
A RAT isn't a replacement for other adversarial techniques — it's the fastest, narrowest one, built to falsify a single assumption rather than stress-test an entire plan. Premortems and persona critiques surface more assumptions at once; a RAT loop decides which one of those is worth testing this week.
Here's how the pieces typically fit together, and when to reach for each one:
| Framework | What It Pressure-Tests | Reach For It When |
|---|---|---|
| Riskiest Assumption Test (this article) | The single highest-stakes, lowest-evidence belief behind a bet | You're about to commit budget or a sprint to an unproven idea |
| Four critics premortem panel | Multiple failure modes at once, from four vantage points | Before a launch or a major roadmap commitment |
| Skeptical engineer critique | Feasibility assumptions hiding inside "we can build this" | The riskiest unknown is technical, not commercial |
| Revenue hawk critique | Business-model and monetization assumptions | The riskiest unknown is "will anyone pay for this" |
| Jobs to Be Done | Whether you understand the job customers hire you for | Your evidence column is thin on customer motivation |
| Customer journey mapping | Where behavioral evidence actually comes from | You need evidence built from what people do, not say |
None of these compete with each other. A premortem or persona critique is often how you populate the assumption map in the first place; the RAT loop is how you decide which item on that map gets tested before the rest of the roadmap does.
Keep the Riskiest Bet Visible Instead of Buried in a Doc
The biggest practical failure of assumption mapping usually isn't the framework — it's that the map lives in a slide or spreadsheet nobody reopens after the kickoff meeting. The riskiest assumption needs to stay visible next to the decision it's attached to, not filed away the moment the meeting ends.
This is the specific gap Prodinja's Journals are designed to close. An assumption you surface in a Journal entry can be tracked against a Situation — the decision or bet it belongs to — so the riskiest one stays attached to the thing it's risking, instead of getting buried in a doc nobody opens again before launch.
None of this replaces judgment. It just makes the judgment call visible enough for someone else to challenge it before the bet is already spent.
Key Takeaways
- Confidence measures familiarity, not accuracy — an assumption feels solid because you've discussed it often, not because it's been tested.
- Score assumptions on two axes: impact if wrong, and current evidence. The high-impact, low-evidence quadrant is your riskiest assumption, and the only one that deserves a test cycle right now.
- Evidence quality matters more than evidence quantity — per The Mom Test, a friendly conversation full of leading questions is closer to noise than data.
- A Riskiest Assumption Test is small, fast, and falsification-first — set the kill threshold before running the test, and run it in days, not sprints.
- Adversarial techniques feed the assumption map — a premortem panel or persona critique is often how the riskiest assumption gets surfaced in the first place.
- The loop repeats — killing or confirming one assumption doesn't end the process; it re-ranks the map and moves to the next riskiest one.
- Visibility beats a one-time exercise — an assumption map nobody reopens after kickoff is barely better than no map at all.
Frequently Asked Questions
What's the difference between a riskiest assumption test and an MVP?
A Minimum Viable Product tests a whole value proposition with real users; a Riskiest Assumption Test targets one specific belief with the smallest experiment that could disprove it. Tristan Kromer proposed RAT specifically because MVP had stretched to mean almost anything, while a RAT stays scoped to a single falsifiable claim.
How many assumptions should a team map before picking one to test?
List every assumption the bet depends on — usually somewhere between five and fifteen for a typical feature or initiative — rather than pre-filtering to the ones that feel important. Filtering early is exactly how the riskiest assumption gets missed, since teams are poor judges of their own blind spots.
What actually counts as evidence on an assumption map?
Evidence is anything that could have proven the assumption false and didn't — a real transaction, a documented behavior, a technical prototype that actually ran. A friendly conversation, a hypothetical "I'd probably use that," or an internal opinion doesn't qualify, no matter how many times it's been repeated in a meeting.
Is a riskiest assumption test the same thing as a premortem?
No. A premortem imagines a future failure and works backward to list what could have caused it, typically surfacing many risks at once; a RAT tests one specific assumption in the present with a real, falsifiable experiment. They're complementary — run a four critics premortem panel to generate the list, then run a RAT loop on whatever tops it.
How fast should a RAT loop actually take?
Days, not weeks — if a test needs more than about a week to design and run, it has usually grown into a small build rather than staying a falsification test. The speed is the point: a RAT that takes as long as building the feature defeats its own purpose.