An OKR smell test is a short, structured audit you run on every draft Key Result before the quarter locks, checking each one against nine recurring failure patterns: disguised to-do lists, vanity metrics, unfalsifiable objectives, too many goals, missing owners, missing baselines, sandbagged targets, output-only measures, and copy-pasted cascades. Catching these before commitment is far cheaper than discovering them at the check-in.

Quick answer: Before locking the quarter, test every draft OKR against nine tells: is it a task, a vanity number, unfalsifiable, one of too many, ownerless, baseline-free, sandbagged, output-only, or copy-pasted from another team? A single "yes" means rewrite — don't ship it as-is.

Why Most OKR Reviews Miss the Real Problems

Most OKR reviews check formatting, not diagnosability — grammar, deadline, whether a number exists — so anti-patterns that look fine sail straight through lock. A smell test asks a different question: can this Key Result actually fail? Would a stranger reading it in week nine, with no context, know whether the quarter succeeded?

That distinction traces back to the framework's own origin. Andy Grove built OKRs at Intel as an evolution of Peter Drucker's Management by Objectives, explicitly to fix MBO's tendency to reward activity over results — a lesson documented in his book High Output Management. John Doerr carried that discipline to Google in 1999 and popularized it in Measure What Matters. The through-line in both: an objective without a falsifiable, ownerable, honestly-scored result isn't an OKR — it's a wish with a due date.

Christina Wodtke, author of Radical Focus, adds the sharpest test for whether a goal is worth locking: could this Key Result not be hit? If every plausible version of the quarter ends in a checkmark, it was never really a goal. That single question kills more anti-patterns than any formatting rule.

Nine specific smells recur across almost every OKR set a PM will review. Group one lives inside the Key Result itself — the wording problem. Group two lives in the structure around it — the goal-setting-process problem. Both groups need a separate pass, because a KR can be perfectly worded and still be the fifth objective too many.

Smells Inside the Key Result

These five show up in the metric line itself — the part everyone reads fastest and audits least.

1. To-Do Lists Wearing KR Clothing

The tell: the Key Result starts with a verb and ends with a deliverable — "Launch the new onboarding flow," "Ship v2 of the pricing page," "Complete the API migration." It describes an activity, not a state of the world that either happened or didn't at a measurable level.

Doerr's own framing makes the distinction explicit: a Key Result should describe an outcome you can score, not a project you can check off. A task either got done or it didn't — there's no 0.3 or 0.7 version of "launched," which is exactly why it can't carry a confidence signal through the quarter.

The fix: rewrite the deliverable as the change it's supposed to cause. "Launch the new onboarding flow" becomes "Increase week-1 activation from 34% to 45%." The launch is the plan; the KR is what the launch was for.

2. Vanity Metrics

The tell: the number moves easily, moves in only one direction, and doesn't force any real trade-off — total signups, page views, app downloads, "number of features shipped." It looks quantitative, which is what makes it easy to wave through review.

Vanity metrics are a close cousin of what the complete guide to killing goal theater in the quarterly cycle calls performative goal-setting: numbers chosen because they're safe to report, not because they're diagnostic of anything. A metric that only ever goes up isn't measuring a decision — it's measuring time passing.

The fix: pair every count metric with a rate or a quality gate. "10,000 signups" becomes "10,000 signups with 30-day retention ≥ 25%" — now growth and quality can't be traded off against each other silently.

3. No Baseline

The tell: the target is a bare number with no visible starting point — "Reach 90% CSAT," "Get NPS to 50." Nobody reviewing it can tell if that's a stretch, a freebie, or a regression dressed up as a goal.

A target without a baseline is unauditable by construction — you cannot tell effort from noise. This is also where a lot of teams discover they never actually captured a real starting point, because the metric was never measured before someone decided to set a goal around it. Building an honest baseline often means walking the customer journey to find where the current experience actually sits today, not where the team assumes it sits.

The fix: every target gets a "from → to." "Reach 90% CSAT" becomes "CSAT from 71% (Q2 baseline) to 80%." If there's no reliable baseline yet, the first-quarter KR is to establish one — that's a legitimate goal, not a cop-out.

4. Sandbagged Targets

The tell: the team scores 0.9–1.0 on this KR every single quarter. It always lands right on target, quarter after quarter, with suspiciously little drama in between.

Google's own internal re:Work guidance on OKRs is explicit that a healthy team average lands around 0.6–0.7 — scoring near 1.0 consistently is treated as a signal the targets were set too low, not that performance was too good. A KR that never surprises anyone isn't stretching anything; it's a target picked backward from what was already going to happen.

The fix: ask the owner, out loud, in the review: "what's the scenario where we miss this?" If nobody can describe one, raise the target until a miss is genuinely possible.

5. Output-Only KRs

The tell: the Key Result counts things produced — features shipped, experiments run, tickets closed, articles published — with no downstream measure of whether any of it changed a customer's behavior or a business number.

This is the output-versus-outcome distinction in its purest form, and it's the one Marty Cagan and the Silicon Valley Product Group return to constantly: shipping is an input to a result, not the result itself. A team can hit every output KR and leave the underlying problem completely untouched.

The fix: for every output KR, ask "output of what, toward what?" and add the missing half. "Ship 12 experiments" becomes "Ship 12 experiments; ≥3 move activation by a statistically significant margin."

The Rewrite, Side by Side

Seeing all five KR-level fixes next to their original wording makes the pattern easier to spot in your own drafts than reading the rules alone.

Anti-patternBefore (fails the smell test)After (passes)
To-do list as KRLaunch the new onboarding flowIncrease week-1 activation from 34% to 45%
Vanity metricGrow signups to 10,00010,000 signups with 30-day retention ≥ 25%
No baselineReach 90% CSATCSAT from 71% (Q2 baseline) to 80%
Sandbagged targetHit 95%+ of sprint commitments, four quarters runningCut lead time from 12 days to 7 days
Output-only KRShip 12 experimentsShip 12 experiments; ≥3 move activation by a statistically significant margin

Notice the shape every fix shares: a bare activity or a bare number turns into a measurable change with a direction and a comparison point. That's the same edit nine times over, just applied to a different smell each time.

Smells in the Goal Structure

These four don't live in any single KR's wording — they show up when you zoom out to the whole objective set.

6. Unfalsifiable Objectives

The tell: the objective is a mood, not a claim — "Delight our customers," "Be the best platform for X," "Build a world-class team." No Key Result underneath it could conceivably prove the objective false.

Locke and Latham's decades of goal-setting research (the foundational organizational-psychology work behind most modern goal frameworks) found specific, measurable goals consistently outperform vague "do your best" ones on actual behavior change. An unfalsifiable objective inherits that weakness at the top of the hierarchy, no matter how tight the KRs underneath it get.

The fix: rewrite the objective as a claim someone could argue with. "Delight our customers" becomes "Make renewal the obvious choice for our highest-value segment" — a claim the KRs can then prove or disprove.

7. Too Many Goals

The tell: the team is carrying five, six, or eight objectives with three-plus Key Results each — eighteen-plus numbers everyone is nominally accountable for, and zero anyone can name from memory in the hallway.

Doerr's own guidance, echoed by most OKR coaches including Felipe Castro, caps a healthy set at three to five objectives with three to four Key Results apiece. Beyond that, OKRs stop functioning as focus and start functioning as a KPI dashboard wearing a quarterly label.

The fix: force a rank-order conversation before lock — not "can we cut," but "which three matter enough that missing them would change the quarter's outcome." Everything else demotes to a tracked metric, not a committed objective.

8. No Owner

The tell: the Key Result reads as a team-wide aspiration — "Improve activation" — with no single name attached who is accountable for moving it, escalating blockers on it, and reporting its status at check-in.

A KR with no owner degrades quietly: everyone assumes someone else is watching it until the quarter ends and nobody was. This is a distinct failure from having no baseline — a KR can be perfectly measurable and still die from diffusion of responsibility.

The fix: name one accountable owner per Key Result, even when execution is genuinely cross-functional. Shared execution and single accountability aren't in conflict — only single accountability survives a bad quarter.

9. Copy-Pasted Cascades

The tell: every team's OKRs are the company objective with the team's name swapped in, restated at a slightly smaller scale, all the way down the org chart — a waterfall dressed up as alignment.

This is the exact failure mode the guide to cascading OKRs without the waterfall is built to break: alignment isn't sameness, it's each team owning a distinct piece of the company's result in the language of their own work. Identical KRs at every level usually means nobody below the top actually thought about how their function contributes.

The fix: require every team's Key Results to be expressed in a metric only that team could plausibly move. If a KR could be copy-pasted onto a different team unchanged, it hasn't been localized yet.

The 15-Minute Pre-Lock Audit

Running all nine checks against every draft OKR is fast if you work from a single table instead of re-deriving the logic each time. Read each Key Result against every row before it's allowed into the locked set.

Anti-patternThe tellOne-line fix
To-do list as KRStarts with a verb, ends with a deliverableRewrite as the measurable change the deliverable causes
Vanity metricOnly moves up, forces no trade-offPair the count with a rate or quality gate
No baselineBare target, no visible starting pointAdd "from → to"; if no baseline exists, make measuring it the KR
Sandbagged targetScores 0.9–1.0 every quarterAsk "what's the miss scenario?" — raise the target until one exists
Output-only KRCounts things shipped, not things changedAdd the downstream outcome the output was supposed to produce
Unfalsifiable objectiveA mood, not a claimRewrite as a claim the KRs could disprove
Too many goals5+ objectives, 15+ total KRsRank-order to 3–5 objectives that would change the quarter
No ownerTeam-wide aspiration, no name attachedAssign one accountable owner per KR
Copy-pasted cascadeIdentical KR at every org levelLocalize to a metric only that team can move

Work the table in two passes, not one. First pass: read every KR alone, cold, and flag anything that trips smells 1–5 (the wording-level failures — they're visible without any org context). Second pass: step back and look at the full objective set together for smells 6–9, since those only show up in relation to the other goals around them.

Bring the complete guide to advanced OKRs into the review if the team is newer to the framework — it covers the scoring mechanics and cadence assumptions this checklist takes for granted. A smell test works fastest as a pre-lock gate, not a post-mortem: cheaper to rewrite a draft KR on Monday than to discover in week nine that it never could have failed.

One more source worth pulling in before lock: objectives grounded in a real customer job, not an internal roadmap item, are naturally harder to write as unfalsifiable or vanity-metric goals. Teams that ground their objective language in Jobs to Be Done research tend to write Key Results tied to a job getting done better, which is inherently more falsifiable than "delight our customers."

Where Prodinja Fits

Two of these nine smells — no baseline and unfalsifiable target — are structural, which means tooling can catch them before a human reviewer has to. Prodinja's Outcome thresholds are designed to do exactly that: setting a threshold requires entering both a current baseline value and a target value before it saves, so a goal with no starting point simply can't be recorded as locked.

That doesn't replace the other seven checks in this list — a threshold field can't tell you an objective is a mood, or that a KR is secretly a to-do list, or that five teams pasted the same target. Those still need a human reading the language, ideally with this table open next to the draft.

But making baseline and target mandatory fields is designed to remove two of the most common failures from the pool before the review conversation even starts — a meaningful head start on a nine-item audit.

Key Takeaways

  • Run the smell test before lock, not at check-in — a rewrite costs minutes pre-commitment and costs a wasted quarter after.
  • Split the audit into two passes: KR-level wording smells (1–5) read cold, objective-set smells (6–9) read in context.
  • A Key Result should be falsifiable — if every plausible version of the quarter ends in a checkmark, per Christina Wodtke's test, it was never a real goal.
  • Sandbagging hides inside a good-looking scorecard — Google's own re:Work guidance treats a 0.6–0.7 average as healthy, not a 0.9–1.0 streak.
  • Output and vanity metrics are the two most common disguises for a KR that doesn't actually measure customer or business change.
  • Fewer, sharper goals beat a long list — three to five objectives someone can name from memory beats eighteen numbers nobody tracks.
  • Baseline and falsifiability are the two smells tooling can catch structurally, which is what Prodinja's Outcome thresholds are built to enforce before a goal locks.

Frequently Asked Questions

How many OKR anti-patterns should I check before locking the quarter?

Nine is a practical, comprehensive floor: five wording-level smells inside each Key Result (to-do lists, vanity metrics, no baseline, sandbagged targets, output-only measures) and four structural smells across the objective set (unfalsifiable objectives, too many goals, no owner, copy-pasted cascades). Fewer checks and real failures slip through; more and the audit stops being fast enough to actually run every quarter.

What's the difference between a vanity metric and an output-only KR?

A vanity metric is a real outcome number that's been chosen because it's easy and always trending up — total signups, page views. An output-only KR isn't measuring an outcome at all; it's counting activity — features shipped, experiments run. Both fail the same test (does this number reflect a real trade-off or customer change?), but the fix differs: pair a vanity metric with a rate, and pair an output KR with the outcome it was meant to produce.

Is scoring 1.0 on every OKR actually a bad thing?

Usually, yes. Consistently landing at or near 1.0 is one of the clearest structural tells that targets were set too low rather than that performance was exceptional — Google's public re:Work OKR guidance treats roughly 0.6–0.7 as the healthy average score for a stretch-oriented team. A team that never misses isn't setting goals; it's forecasting outcomes it already expected.

How do I fix an OKR with no owner without creating a bottleneck?

Assign one accountable name to the Key Result even when the work is genuinely cross-functional — accountability for reporting status and escalating blockers isn't the same as doing all the execution alone. The owner's job is to know the number's status at any point in the quarter and pull the right people in when it's off track, not to personally move every input to it.

Can OKR anti-patterns be caught by software, or does it require a human review?

Some can be caught structurally — a tool can require a baseline value and a target value before a goal saves, which closes off the no-baseline and unfalsifiable-target smells by construction. The other seven — disguised to-do lists, vanity metrics, sandbagging, output-only framing, unfalsifiable objectives, too many goals, and copy-pasted cascades — are language and judgment calls that still need a human reader working through a checklist like this one before the cycle locks.