There is no universal number of user interviews that guarantees good decisions — five, twelve, and thirty are all defensible depending on your segments and stakes. The real signal is thematic saturation: the point where new interviews stop surfacing new themes and start confirming what you already heard.

Quick Answer: Stop counting interviews and start counting new themes. When 2-3 consecutive interviews add nothing new within a single user segment, you've likely reached saturation for that segment. Multiple segments each need their own saturation point.

Why "Five Users" Became a Rule of Thumb, Not a Law

The "five users" heuristic comes from a specific, narrow research context — usability testing on a single interface, within one user segment. It was never meant to answer "how many interviews for discovery."

Jakob Nielsen and Tom Landauer's 1993 analysis of usability studies found that the first five users typically uncover roughly 80% of usability problems in a single, homogeneous user group testing a specific interface. That math holds because usability problems are a bounded, discoverable set — there's a finite number of ways people get confused by one screen. It breaks down the moment you generalize it to discovery research, where you're not counting bugs, you're mapping an open-ended landscape of needs, motivations, and workflows across possibly multiple, different kinds of users.

The rule got flattened in translation. Teams heard "five users is enough" and applied it to everything from onboarding usability tests to full-blown market discovery, without checking whether the underlying assumption — one homogeneous group, one bounded interface — still applied. It usually doesn't.

What Nielsen's Model Actually Assumes

Three conditions have to hold for "five is enough" to be true, and discovery research routinely violates at least one:

  1. A single user segment. Mixing power users, casual users, and admins collapses distinct populations into one average, which hides exactly the variation you're trying to find.
  2. A bounded problem space. Usability defects are finite; motivations, jobs-to-be-done, and workflow variations are not bounded the same way.
  3. A specific artifact under test. You're evaluating reactions to something concrete, not open-ended exploratory questions about problems, context, or unmet needs.

If you're running generative discovery interviews — asking about jobs, frustrations, and workarounds rather than testing a prototype — treat "five" as a floor per segment, not a target.

What Thematic Saturation Actually Looks Like

Thematic saturation is the point in a qualitative study where additional interviews stop producing new codes, themes, or insights relevant to your research question — not the point where people start repeating exact quotes. It's a pattern-level signal, not a word-level one.

Greg Guest, Arwen Bunce, and Laura Johnson's widely-cited 2006 study on data saturation found that roughly 92% of all themes emerging across sixty interviews had already appeared within the first twelve — and most of the foundational, high-frequency themes showed up within the first six. That's directionally consistent with what most experienced researchers report: diminishing returns set in early, but the tail of rare-but-important themes keeps trickling in much longer.

The Saturation Grid

Track saturation the way you'd track a code review — as a running tally, not a gut feeling:

Interview #New themes surfacedCumulative unique themesSignal
1-3High (5-8 each)Rapidly climbingToo early to conclude anything
4-6Moderate (2-4 each)Growth slowingStill exploratory
7-9Low (0-2 each)Near-flatApproaching saturation
10-12Near-zero (0-1 each)FlatLikely saturated for this segment

The mechanism that makes this useful is comparison, not counting. You need a consistent way to log a theme against the interview that produced it, so you can literally see the curve flatten. This is close to what Journals in Prodinja is designed for: tag recurring themes across sessions as you capture them, so the flattening curve becomes visible instead of something you have to reconstruct from memory or scattered notes after the fact.

Three Practical Markers of Saturation

  • Redundancy, not repetition. You'll hear different words for the same underlying problem — that's still saturation. Watch for the idea, not the phrasing.
  • Confidence in prediction. You can accurately guess what the next participant will say before they say it, and you're usually right.
  • Diminishing surprise. Early interviews reliably produce a "huh, didn't expect that" moment. When several interviews in a row produce none, that's your signal, not a fixed count.

Why Segmentation Multiplies Your Interview Count

Every distinct user segment resets your saturation clock — a mistake teams make constantly is running twelve interviews across three segments and treating that as "twelve interviews' worth" of saturation, when it's really four per segment, well short of typical saturation thresholds.

Segments are not demographic buckets — they're behavioral or contextual differences that plausibly change the answer to your research question. A B2B SaaS product might need separate saturation for admins versus end users versus economic buyers, because their jobs, incentives, and frustrations genuinely differ. A consumer app might not need segmentation by industry at all, but might badly need it by usage frequency.

A Simple Segmentation Test

Ask this before you plan interview count: would a finding from segment A actually surprise or contradict what you'd expect from segment B? If yes, they're separate segments and each needs its own saturation runway. If no — if you're just imagining a difference that doesn't change behavior — collapse them and save the interviews.

Segmentation approachInterviews needed (rule of thumb)Risk if skipped
Single segment, narrow question6-12Low — usability-style questions saturate fast
2-3 meaningfully distinct segments8-15 per segmentMedium — missing a segment's needs entirely
Cross-functional (buyer, user, admin)10+ per roleHigh — solving the wrong person's problem
New/unfamiliar market or persona15-20+, unclear until you're in itVery high — no prior pattern to anchor against

This is also where the product discovery complete guide is worth reading in full — segmentation strategy sits upstream of almost every interview-count decision, and getting it wrong upstream makes every downstream number meaningless.

Watch the Interaction With Question Quality

Saturation only means something if the underlying interviews are actually good. A set of interviews dominated by leading or closed questions will "saturate" fast — because you're hearing your own assumptions echoed back, not new information. Before trusting a flattening curve, sanity-check that your interview guide is actually built to surface disagreement; the distinction between open, leading, and closed interview questions is the difference between real saturation and a false plateau.

How Continuous Discovery Reframes the Question Entirely

The "how many interviews" question assumes a project has a start and an end — but most modern product teams run discovery as an ongoing habit, which makes total sample size the wrong metric to optimize for in the first place.

Teresa Torres, who popularized the continuous discovery model, argues for a standing cadence — commonly at least one customer touchpoint per week, sustained indefinitely — rather than a single research sprint sized to hit a saturation target. Under this model, saturation isn't something you calculate once and stop; it's something you continuously re-check as your product, market, and users change, because a "saturated" understanding from eight months ago can quietly go stale.

Why Cadence Beats a One-Time Count

  1. Products and users drift. A saturated theme set from a pre-launch study won't reflect how usage patterns change six months post-launch.
  2. New segments emerge as you grow. Enterprise deals bring admin personas you never interviewed; international expansion brings context you've never tested against.
  3. Weekly cadence keeps the cost low and the risk of stale assumptions lower. Instead of a one-time 40-interview push, a small steady stream means you're never more than a week or two from your last real signal.

If you've never run a standing research habit, the weekly discovery habit built on two interviews is a concrete, low-friction way to start one without requiring a research team or a large time budget.

Reframing the Metric

Old framingContinuous discovery framing
"How many interviews before we ship?""Have we talked to a user this week?"
One-time saturation study, then silenceOngoing low-volume cadence, revisited regularly
Sample size decided upfrontSample size emerges from watching the theme curve
Discovery ends when research endsDiscovery is a standing team habit

This doesn't mean saturation stops mattering — it means you apply the saturation logic from the sections above continuously, in small batches, instead of trying to front-load it into one large study that goes stale the moment it's finished.

When to Stop, Segment, or Keep Going: A Decision Framework

Use three questions in sequence rather than a fixed number, because the right call depends on what's actually flattening — or not — in your theme log.

First, check the curve. If the last 2-3 interviews within a single segment produced zero new themes relevant to your research question, you've likely hit saturation for that segment — stop there and move to synthesis.

Second, check for hidden segments. If themes keep appearing but seem to cluster by a variable you haven't controlled for (role, tenure, company size, usage frequency), you're not unsaturated — you're under-segmented. Split your sample and re-run the saturation check per segment.

Third, check the stakes. A low-stakes UI tweak can tolerate a thinner saturation bar than a new product line or a pricing change — weight your interview investment to the size of the bet, not a fixed number regardless of context.

A Compact Decision Table

SituationAction
New themes have stopped appearing for 2-3 interviews in a rowStop this segment, move to synthesis
New themes correlate with a variable you didn't segment onSplit the sample, restart saturation tracking per segment
High-stakes decision, thin saturation so farKeep going — invest proportional to the bet size
Low-stakes decision, moderate saturation reachedStop — diminishing returns don't justify more interviews
It's been weeks since your last user conversationIgnore the count entirely — go talk to someone this week

Once you've reached saturation on the "what" — the themes and frustrations — mapping them onto a jobs-to-be-done framework or plotting them against a customer journey is usually the more valuable next move than squeezing out more interviews on the same theme set. Saturation tells you when to stop collecting; it doesn't do the synthesis for you — an opportunity solution tree built for shipping teams is one structured way to turn a saturated theme list into prioritized bets.

Key Takeaways

  • "Five users" is a usability-testing heuristic, not a discovery research law — it assumes one homogeneous segment and one bounded interface, conditions discovery interviews rarely meet.
  • Thematic saturation, not interview count, is the real stopping signal — track new themes per interview and stop when 2-3 in a row add nothing new.
  • Every distinct user segment resets the saturation clock — twelve interviews across three segments is four per segment, not twelve interviews' worth of saturation.
  • Question quality determines whether a plateau is real — leading or closed questions produce a false, premature saturation signal.
  • Continuous discovery reframes "how many" into "how often" — a standing weekly cadence keeps understanding current instead of letting a one-time study go stale.
  • Weight interview investment to the stakes of the decision — a minor UI tweak and a new product line don't deserve the same sample size.
  • Tagging themes as you go, rather than reconstructing them afterward, is what actually makes a saturation curve visible — without a running log, "saturation" is just a feeling.

Frequently Asked Questions

How many user interviews do I need for a new feature?

For a single, well-defined user segment, most teams see thematic saturation somewhere between 8 and 15 interviews. That number roughly doubles or triples if the feature affects multiple distinct segments, since each one needs its own saturation check.

Is 5 user interviews really enough?

Five interviews is a usability-testing benchmark from Nielsen and Landauer's research, not a discovery research standard — it assumes one homogeneous user group testing one specific interface. For open-ended discovery interviews about needs and workflows, treat five as an early checkpoint, not a stopping point.

How do I know when I've reached saturation in qualitative research?

You've likely reached saturation when 2-3 consecutive interviews within the same segment fail to surface any new theme relevant to your research question. Track themes per interview in a running log so you can see the curve flatten rather than relying on impression alone.

Does sample size matter more than interview quality?

No — a larger sample of leading or closed-question interviews will produce a false, premature saturation signal, because you're mostly hearing your assumptions echoed back. Interview quality determines whether a flattening theme curve reflects reality or just weak question design.

Should I keep doing interviews after launch?

Yes — continuous discovery treats research as an ongoing weekly habit rather than a one-time pre-launch study, because user needs and product usage both drift over time. A saturated understanding from months ago can quietly go stale without a standing cadence to catch the drift.