Tying bonuses to OKR scores does not raise ambition — it lowers it. The moment a scorecard number determines pay, people do what any rational actor would: negotiate softer targets, avoid genuine stretch goals, and stop surfacing bad news early. Decoupling OKRs from compensation is what keeps goals honest and ambitious again.

Quick answer: When OKR scores drive bonuses, people set targets they know they can beat — a well-documented pattern economists call the ratchet effect. Score OKRs for direction and learning; evaluate compensation on contribution, judgment, and growth instead, using a separate rubric entirely.

What Actually Happens the Moment OKR Scores Set Bonuses

Once an OKR score feeds a bonus formula, target-setting stops being a planning exercise and becomes a negotiation. People sandbag: they lowball the baseline, pad the timeline, and quietly steer away from the objective most likely to embarrass them if missed. Ambition turns from a virtue into a liability.

This isn't a character flaw — it's a predictable response to a badly designed system. Economist Charles Goodhart's famous observation, now shorthand as Goodhart's Law, applies directly the moment money enters the picture.

When a measure becomes a target, it ceases to be a good measure. Attach money to a 0–1.0 OKR score and the score stops measuring progress — it starts measuring how well someone can negotiate a beatable number.

The research on this is not new or fringe. In "Goals Gone Wild," a widely cited 2009 paper in Academy of Management Perspectives, Lisa Ordóñez, Maurice Schweitzer, Adam Galinsky, and Max Bazerman reviewed decades of goal-setting research and found that goals tied to rewards reliably produce narrower focus, more risk-taking, and more unethical shortcuts than goals set without a financial trigger attached.

Behavioral economists have documented the same corrosion from the motivation side. Uri Gneezy and colleagues, reviewing incentive experiments in the Journal of Economic Perspectives, found that explicit financial incentives can crowd out intrinsic motivation — people who were previously stretching themselves for the work itself start doing the minimum that clears the bar once a bonus is attached to it.

Three behaviors show up almost every time OKRs and comp get linked:

  • Sandbagging at target-setting. A team that could plausibly hit a 40% improvement proposes 15%, because a missed stretch goal now costs real money.
  • Goal narrowing. People over-invest in the one key result being measured for pay and quietly neglect everything adjacent to it — even work that matters more.
  • Delayed bad news. A key result that's clearly off-track in week three gets reported as "on track" in the mid-quarter check-in, because admitting trouble early feels like admitting a pay cut.
BehaviorOKRs used as a planning toolOKRs used as a comp scorecard
Target-settingTeams propose the honest stretch numberTeams propose the safely beatable number
Mid-cycle reportingOff-track KRs get flagged early for helpOff-track KRs get quietly reframed as "on track"
Cross-team asksPeople pursue the highest-leverage workPeople protect only the KR tied to their bonus
FailureA missed KR triggers a retro and a new betA missed KR triggers a dispute about the payout
Best next quarterTargets get more ambitious as trust buildsTargets converge toward whatever was safe last time

This is exactly the mechanic behind what our guide to killing goal theater in the quarterly cycle describes: a check-in ritual that looks like honest progress reporting but is quietly optimized for how it will read at comp time, not for whether the team is learning anything.

It's also what makes cascading OKRs without creating a waterfall so much harder in practice. A manager who negotiates a soft number to protect their team's bonus hands every team below them a softened target too, and the corruption compounds at every layer down the org chart.

OKRs Point the Direction; Reviews Judge the Person — Two Different Jobs

OKRs exist to align a group of people around what matters most this quarter and to make progress visible while there's still time to act on it. Performance reviews exist to judge an individual's contribution, judgment, and growth over a longer arc. Collapsing both into a single number destroys what each was built to do.

An OKR score is, by design, a team-level, time-boxed, partly-luck-dependent number. A key result can miss because a dependency team slipped, a market shifted, or a competitor moved first — none of which says anything about whether the person who owned it did excellent work. Our complete guide to OKRs covers this distinction between the objective layer and an individual's actual judgment in more depth.

A missed key result tells you the bet didn't land. It does not tell you whether the person made a good bet, executed it well, and learned fast when it started to go sideways.

This is precisely why John Doerr's Measure What Matters, the book that popularized OKRs outside Intel and Google, is explicit that OKRs were never designed as the input to a comp decision. Google's own internal OKR guidance — documented publicly through its re:Work resource — treats OKR scores as a grading exercise for the goal, not a rating of the person who owned it, and keeps the two processes on separate calendars with separate owners.

That separation only works if you also get the OKR itself right. A key result written as an output — "ship the redesign" — is trivially satisfiable regardless of impact, which makes it a worse scorecard for pay decisions, not a better one. Our piece on the outcome vs. output distinction in OKRs is the companion read here: a badly written KR corrupts a comp decision even faster than a well-written one, because it can be gamed with zero real progress at all.

The direction/judgment split, in one sentence

OKRs answer "are we moving toward the right thing." Performance reviews answer "how well is this person operating." Those are different questions, asked on different timelines, and answered with different evidence — treating them as the same question is the root error underneath every symptom in the section above.

If Not the OKR Score, Then What? A Fairer Way to Evaluate People

Evaluate people on four inputs that have nothing to do with the raw OKR percentage: the quality of judgment behind their bets, the caliber of their execution, how they influenced people outside their direct reporting line, and evidence of growth since the last cycle. None of these require reducing a quarter to a single 0–1.0 number.

This is the manager's real objection, and it's a fair one: "If I can't point to the OKR score, how do I justify a rating without it looking arbitrary?" The answer is a documented rubric, evaluated with real evidence, the same way a hiring panel evaluates a candidate — not a vaguer process, a differently-structured one.

Evaluation dimensionThe question a manager asksWhat counts as evidence
JudgmentDid they bet on the right problem, and can they explain why?Did they ground the objective in a real customer job to be done, or a specific friction point on the customer journey, rather than a guess or a leadership request taken at face value?
ExecutionDid the work they controlled get done well, on the things that were theirs to control?Shipped quality, rework rate, how they handled a dependency slipping
InfluenceDid they move people or teams outside their formal authority?Cross-team asks landed, stakeholders aligned without escalation
GrowthAre they visibly better at the job than they were last cycle?New skill demonstrated, harder problem taken on, feedback acted on

Notice that "judgment" explicitly rewards the discovery work behind the goal, not just the number the goal produced. A PM who used a rigorous method like Jobs to Be Done to find the real problem, or who traced a customer journey to the actual friction point instead of guessing, made a better bet — even in a quarter where the resulting key result came up short.

This isn't a new idea invented to solve the OKR problem — it's borrowed from companies that separated these processes long before OKRs got popular:

  1. Netflix's "Keeper Test" — documented in the company's widely circulated culture memo — asks managers whether they'd fight to keep someone, a judgment call built on the whole picture of a person's work, not a single quarterly metric.
  2. Adobe's "Check-In" system, which replaced Adobe's annual review and forced-ranking process in 2012, moved to ongoing manager conversations about expectations and growth, explicitly separate from any single scored target.
  3. Christina Wodtke's Radical Focus recommends running the goal-setting ritual and the people-development ritual on entirely different cadences, precisely so neither one contaminates the other.

How to Decouple Without Looking Like You're Going Soft on Accountability

Decoupling OKRs from comp is not the same as removing accountability — it's moving accountability to evidence that's actually fair to judge someone on. The switch takes five concrete moves: announce it explicitly, separate the tools, change what check-ins are for, build a real comp rubric, and calibrate across managers so it doesn't become manager-dependent luck.

  1. Say it out loud, in writing. If people have quietly believed OKR percentage drives their bonus, an unannounced policy change reads as a broken promise. State plainly: "OKR scores measure our bets, not your rating."
  2. Put the two processes in different systems. If the comp conversation and the OKR tracker live in the same tool, it's psychologically hard to fully unlink them. Score OKRs somewhere; evaluate people somewhere else, on a different rhythm.
  3. Turn check-ins into learning conversations, not status reports. Ask "what did we learn and what should we bet differently next cycle," not "are we going to hit the number" — the second question is the one that produces sandbagging.
  4. Build a comp rubric with real dimensions, like the judgment/execution/influence/growth table above, and require managers to cite specific evidence, not a percentage, when they write a rating.
  5. Calibrate across managers before ratings go final. A rubric without calibration just relocates the unfairness from "which team had an easier OKR" to "which manager rates generously" — a calibration session catches both.

Give managers a genuinely separate lane for growth

The hardest part of step 2 is practical, not philosophical: most teams don't have a place to track judgment, execution, and growth evidence that isn't the same spreadsheet or tool tracking the OKRs themselves.

Prodinja, which currently ships as an interactive prototype, is built around that separation rather than papering over it. Its Growth competencies and Leadership Suite give a manager a distinct, development-focused lane for logging judgment, contribution, and skill evidence over time — deliberately apart from the OKR planning surface, so a manager isn't tempted to reach for the nearest number just because it's the one already open on screen.

Keeping the review evidence and the OKR tracker in physically separate places is a small structural nudge toward the same discipline this whole article is arguing for: judge the person on the person's evidence, not the goal's outcome.

Objections Leaders Raise — and What's Actually True

Every leader who considers decoupling hears the same pushback within the first meeting: won't people just coast without a number attached to pay? The honest answer is that a badly designed link to comp doesn't prevent coasting — it just hides it behind a negotiated-down target instead.

ObjectionWhat's actually true
"Without OKR-linked bonuses, nobody will push hard."Ambition drops because of the link, per Ordóñez et al.'s findings — decoupling tends to raise the targets people are willing to propose, not lower effort.
"This removes accountability."It relocates accountability to judgment, execution, and growth evidence — a rubric, not a vibe — which is harder to game than a single percentage.
"It's more work for managers."It's different work, concentrated at calibration time instead of spread across every anxious mid-quarter check-in where people are managing the score, not the goal.
"Google and Intel are huge — this won't scale to us."The practice predates their scale: Andy Grove ran OKRs this way at Intel from the start, specifically to keep target-setting honest under a much smaller company.
"Won't top performers feel under-rewarded?"Top performers are usually the ones most frustrated by sandbagging culture — they're the ones capable of hitting real stretch targets and penalized for it under a comp-linked system.

The pattern in every row is the same: the objection assumes comp-linked OKRs are the thing producing rigor, when the evidence says they're the thing quietly eroding it. Decoupling doesn't remove the rigor — it moves the rigor to where it can actually be assessed honestly.

Key Takeaways

  • Tying OKR scores to bonuses produces sandbagging, not ambition — people propose targets they know they can beat, a pattern researchers have documented since at least Ordóñez et al.'s 2009 "Goals Gone Wild."
  • Goodhart's Law explains the mechanism: attach money to a measure and the measure stops reflecting reality — it starts reflecting how well someone can negotiate the number.
  • OKRs and performance reviews answer different questions. OKRs judge the bet; reviews judge the person's judgment, execution, influence, and growth.
  • A missed key result is not proof of poor performance — it can be the result of a market shift, a dependency, or an honest, well-reasoned bet that didn't land.
  • Build a comp rubric with real dimensions — judgment, execution, influence, growth — evaluated with cited evidence, not a percentage pulled from the OKR tracker.
  • Calibrate across managers, or a rubric just relocates unfairness from "whose OKR was harder" to "whose manager rates generously."
  • Keep the tools physically separate. If comp evidence and OKR tracking live in the same system, the unlinking is nominal, not real.

Frequently Asked Questions

Should OKRs be tied to performance reviews at all?

Not directly — OKRs should inform a review as context, not function as its score. A missed key result is useful evidence about market conditions and prioritization, but treating the raw percentage as the rating collapses two different questions (was the bet good, was the person good) into one number that answers neither well.

How do you evaluate performance without using OKR scores?

Use a rubric built on judgment, execution, influence, and growth, evaluated with specific cited evidence — what problem they chose to solve and why, how well they executed what was in their control, who they moved outside their formal authority, and what they're visibly better at than last cycle. This is closer to how a hiring panel evaluates a candidate than how a scoreboard is read.

Does decoupling OKRs from compensation mean removing accountability?

No — it relocates accountability from a single, gameable percentage to a documented, evidence-based rubric that's actually harder to fake. Teams that decouple typically report more rigorous performance conversations, not fewer, because managers have to cite specific evidence instead of pointing at a number.

What do companies like Google actually do about OKRs and bonuses?

Google's own re:Work OKR guidance treats OKR scores as a grading exercise for the goal itself, run on a separate calendar from performance ratings, which are informed by broader judgment- and peer-feedback-based evidence. The scores stay visible and honest specifically because nobody's pay depends on inflating them.

What if leadership insists on linking OKR scores to bonuses anyway?

Push for a lagging, blended approach instead of a direct formula: let OKR outcomes inform a manager's narrative during calibration, alongside the judgment/execution/growth rubric, rather than feeding a percentage straight into a payout calculation. A blended, evidence-weighted view is far harder to game than a single number, even if full decoupling isn't yet politically possible.