An experience "actually works" when it clears four measurable bars: users complete the task (task success rate), do it efficiently (time-on-task), do it without mistakes (error rate), and rate it favorably (SUS score). Track those four plus the HEART framework for ongoing signals, and you replace "it feels smoother" with a number a VP can defend in a budget review.
Quick Answer: Measure UX with task success rate, time-on-task, error rate, and the System Usability Scale (SUS) from a moderated benchmark study, then monitor Happiness, Engagement, Adoption, Retention, and Task success (HEART) in product analytics between studies. Set numeric thresholds for each in your acceptance criteria before a release ships.
Why "It Feels Better" Stops Working in a Budget Review
Subjective UX feedback dies the moment a finance stakeholder asks for a number, because "feels more intuitive" has no unit and no baseline to compare against. Quantified usability data survives that question because it's comparable across releases, teams, and competitors. This is the entire case for measurement discipline.
The shift isn't cosmetic — it changes what a PM can promise. "We think the new onboarding is cleaner" is an opinion nobody has to act on. "Task success rate on account setup moved from 61% to 84%, and time-on-task dropped 40 seconds" is a claim a stakeholder can weigh against engineering cost.
Jakob Nielsen and Tom Landauer's research, published through the Nielsen Norman Group, has long argued that usability problems are found reliably with structured measurement paired with a handful of test sessions — not intuition alone. That combination, not raw sample size, is what makes small-team UX research defensible.
The Credibility Gap PMs Actually Face
Design teams often already know where friction sits; the gap is translating that knowledge into a format that survives a roadmap prioritization meeting. Numbers travel across functions in a way narrative descriptions don't.
- Engineering wants a defect-style severity, not an adjective.
- Finance wants a before/after delta tied to a business metric.
- Sales/CS wants to know if the number moves churn or ticket volume.
Quant UX metrics are the shared language that satisfies all three audiences from one study.
The Four Core Metrics: Task Success, Time-on-Task, Error Rate, SUS
The foundational quartet — task success rate, time-on-task, error rate, and SUS — comes from usability testing methodology and answers, respectively: did they finish, how long did it take, how many mistakes happened, and how did it feel. Run all four together in the same session for a complete picture; any one alone is misleading.
| Metric | What it measures | Typical target | Common pitfall |
|---|---|---|---|
| Task success rate | % of users who complete a defined task unassisted | 78%+ is considered a reasonable benchmark in NN/g's aggregated usability data | Vague completion definition ("close enough" counted as success) |
| Time-on-task | Seconds/minutes from task start to completion | Compare only against your own baseline, not industry averages | Including think-aloud narration time in a moderated session |
| Error rate | Number of wrong actions per task attempt, successful or not | Trending toward zero for critical-path flows | Conflating slips (misclicks) with mistakes (wrong mental model) |
SUS score | 10-item, 5-point Likert survey converted to a 0-100 scale | 68 is the norm-referenced average per Brooke's original SUS research | Treating a single respondent's score as statistically meaningful |
Task Success Rate: The Binary That Matters Most
Task success rate is the percentage of test participants who complete a defined task without unrecoverable failure. Define "success" before the session starts, in writing, or every observer will silently apply a different bar.
- Write the task scenario in user language, not feature language ("Find last month's invoice," not "Navigate to Billing > History").
- Decide up front whether partial completion with help counts as success, partial success, or failure.
- Record success as binary per participant, then aggregate to a percentage across the session sample.
A completion rate below 70-80% on a core flow, per NN/g's commonly cited threshold, generally signals a usability problem worth fixing before launch rather than after.
Time-on-Task and Error Rate: Efficiency and Friction
Time-on-task is elapsed time from a defined start action to task completion; error rate counts wrong turns per attempt, whether or not the task ultimately succeeds. Together they distinguish "worked but painful" from "failed outright" — two very different engineering priorities.
- Time-on-task only means something against a prior baseline or a competitor's equivalent flow — an absolute number in isolation tells you nothing.
- Error rate should separate slips (accidental misclicks, fixable with better affordance) from mistakes (wrong mental model, fixable with better information architecture) — see the guide on design literacy fundamentals for PMs for the vocabulary to bring that distinction into a design critique.
- A high error rate with a high success rate still matters: it predicts support-ticket volume even when users eventually muddle through.
SUS: The Standardized Feeling Score
SUS is a validated 10-item questionnaire, alternating positive and negative statements, scored on a 0-100 scale after a specific weighting formula. John Brooke published it in 1986 at Digital Equipment Corporation specifically so usability could be compared across products and over time using one instrument.
Per Brooke's original calibration and later replication work by Jeff Sauro at MeasuringU, a
SUSscore around 68 sits at the population average; above 80 is generally considered excellent, and below 51 correlates with real adoption risk.
Administer SUS identically every time — same wording, same 5-point scale, same participant count floor of roughly 12-15 respondents — or comparisons across releases become invalid.
HEART: Tracking UX Continuously Between Studies
Google's HEART framework (Happiness, Engagement, Adoption, Retention, Task success) gives PMs a way to monitor UX quality in product analytics continuously, filling the gap between periodic benchmark studies. Rubin and Chisnell's research on usability testing methodology, and Google's own published HEART work, both frame it as complementary to lab metrics, not a replacement for them.
| HEART dimension | Example signal | Data source |
|---|---|---|
| Happiness | SUS score, CSAT, app-store rating | Survey, benchmark study |
| Engagement | Sessions per week, features used per session | Product analytics |
| Adoption | % of eligible users who try a new feature in 30 days | Product analytics |
| Retention | % of users still active at day 30/60/90 | Product analytics |
| Task success | Completion rate, time-on-task, error rate | Usability testing, funnel analytics |
Map each HEART dimension to Goals-Signals-Metrics before instrumenting anything — a metric with no clear goal behind it tends to get gamed or ignored within two quarters. This is the same discipline that underlies good jobs-to-be-done thinking: define the outcome first, then find the proxy that reflects it.
Where HEART and the Core Four Overlap — and Where They Don't
HEART's task-success dimension literally reuses the completion-rate metric from the lab quartet, which is intentional: it's the connective tissue between one-time studies and always-on telemetry. Happiness, Engagement, Adoption, and Retention, by contrast, need product analytics infrastructure the four core metrics don't.
Run the core four in a moderated study before a launch; watch HEART dashboards after launch to catch regressions the lab study's small sample couldn't. Neither substitutes for the other.
How to Run a Lightweight UX Benchmark Study
A lightweight benchmark needs a defined task list, 5-8 participants per segment, a moderator, and a spreadsheet to log success/time/errors/SUS — no dedicated research team required. Jakob Nielsen's widely cited finding that five users surface roughly 85% of usability problems in a given flow is what makes this tractable for a PM working solo.
- Define 3-5 representative tasks tied to the flow you're evaluating, written as user goals, not UI steps.
- Recruit 5-8 participants per user segment — fewer if budget-constrained, more if the flow serves meaningfully different personas.
- Moderate sessions using think-aloud protocol, logging start/end timestamps, every error, and task outcome live.
- Administer
SUSimmediately after the session, before moving to any follow-up interview questions. - Aggregate and baseline — this first run becomes the number every future release gets compared against.
- Re-run after the fix ships, using the identical tasks and wording, to measure the delta.
Before you script tasks, a quick primer on how cognitive load causes a simple screen to overwhelm users helps you write scenarios that actually surface friction rather than test rote memorization of your own UI.
Turning the Benchmark Into Acceptance Criteria
Numeric UX targets belong in acceptance criteria the same way performance budgets do — as a release gate, not a retrospective nice-to-have. A story that says "improve onboarding" is unfalsifiable; one that says "task success rate ≥ 80%, SUS ≥ 70, error rate ≤ 1 per attempt" is testable.
- Write the target before the design is finalized, using the prior benchmark as the floor.
- Treat a missed target as a blocking defect in the same triage lane as a functional bug, not a backlog item.
- Reference the metric, not the design, in the acceptance criteria — designs will iterate, the bar shouldn't move to match whatever shipped.
This is also where a structured design critique earns its keep: a PM can push back on a design using the missed number, not a personal aesthetic opinion, which keeps the conversation collaborative rather than adversarial.
Where the Emotion Curve Fits Alongside Hard Numbers
Quant metrics tell you how much an experience underperforms; a qualitative emotion curve across the same journey tells you where the drop happens and roughly why. Neither view alone is complete — a low SUS score without a map of the journey just tells you something is wrong, not what.
Prodinja's Customer Journey studio is built around exactly this pairing: it walks you through mapping an emotion curve across journey stages so you can see where sentiment dips, and lets you sit hard metrics like task success rate or time-on-task against the same stages to confirm the dip is real and quantify its size. Used together, the emotion curve nominates the suspect stage and the benchmark study convicts it with a number. For the fuller picture of journey mapping as a discipline, see the guide on customer journey mapping.
That pairing is also what makes a UX investment case land with a skeptical stakeholder: "sentiment drops at checkout step 3" is a hypothesis; "sentiment drops at checkout step 3, and task success there is 58% against an 82% average elsewhere" is a diagnosis.
Key Takeaways
- Task success rate, time-on-task, error rate, and
SUSare the four core metrics that turn "it feels better" into a defensible, comparable number. SUSscores around 68 are average, per Brooke's original research and Sauro's replication work — treat that as your floor, not your ceiling.- Five to eight participants per segment, following Nielsen's usability research, is enough to surface most major usability problems in a lightweight benchmark.
HEART(Happiness, Engagement, Adoption, Retention, Task success) extends measurement into continuous product analytics between benchmark studies.- Numeric UX targets belong in acceptance criteria, set before design finalizes and treated as release-blocking, not retrospective commentary.
- Pairing a qualitative emotion curve with hard metrics — the approach Prodinja's Customer Journey studio is built around — shows both where an experience underperforms and how much.
Frequently Asked Questions
What is a good SUS score for a product?
A SUS score around 68 is the population average per Brooke's original calibration and later research by Jeff Sauro; scores above 80 are generally considered excellent, and anything below 51 correlates with real adoption risk. Always compare against your own product's prior benchmark, not just the population average.
How many users do you need for a UX benchmark study?
Five to eight participants per user segment is typically enough, based on Jakob Nielsen's widely cited finding that five users surface roughly 85% of usability problems in a given flow. Add more participants only if the flow serves meaningfully distinct personas that need separate analysis.
What's the difference between task success rate and HEART's task-success metric?
They're the same underlying measurement — the percentage of users who complete a defined task — but used in different contexts. Task success rate typically comes from a moderated lab study; HEART's task-success dimension usually comes from funnel analytics tracking the same completion event continuously in production.
How do you set UX targets in acceptance criteria?
Run a benchmark study first to establish a baseline, then write the target as a specific number — for example, "task success rate ≥ 80%, SUS ≥ 70" — directly into the acceptance criteria before design finalizes. Treat a missed target as a blocking defect, the same severity tier as a functional bug.
Can you measure UX quality without a dedicated research team?
Yes — a lightweight benchmark study with 5-8 participants, a spreadsheet, and a moderator following think-aloud protocol captures task success rate, time-on-task, error rate, and SUS without specialized tooling. The main requirement is defining tasks and a success threshold before sessions begin, not headcount.