Sandbagging happens when the same OKR number is used to set a stretch target and to judge someone's performance, so people quietly protect themselves by picking a number they know they'll hit. The fix isn't smarter target-setting templates — it's removing the incentive to lie, then rebuilding the calibration muscle that fear wore down.
Quick answer: Sandbagging is caused by coupling OKR attainment to consequences — ratings, pay, headcount. The tell is a Key Result that sits suspiciously close to current run-rate. The fix is decoupling scoring from performance review, making the starting baseline visible to everyone, and renegotiating the target with a shared-risk floor both sides can live with.
Every leader who has run more than two OKR cycles has felt this without naming it: the numbers keep landing at 95-100% attainment, the quarterly review meetings are suspiciously calm, and yet the business isn't moving any faster than it was a year ago. That's not a coincidence. It's what happens when goals stop being instruments of ambition and start being instruments of self-protection. Understanding why requires starting with the incentive, not the target.
Why Leaders Sandbag: The Incentive Root Cause
People set safe targets when missing an ambitious one carries a personal cost and hitting a safe one doesn't. This is a rational response to a badly designed system, not a character flaw — and it traces directly back to whatever your organization has quietly attached to the OKR number, whether or not that attachment is official policy.
The mechanism is old and well documented. Goodhart's Law, named for economist Charles Goodhart, states that once a measure becomes a target, it stops being a reliable measure — because people optimize for the number rather than the underlying reality it was meant to represent. OKRs are exceptionally vulnerable to this because the number is both the goal and, in most companies, an input to how someone is judged.
Trace the causal chain and it's short:
- A Key Result gets tied to a performance rating, formally ("hit 90% of KRs to get a 'meets expectations'") or informally (a manager who remembers who missed their numbers last cycle).
- Missing the target now has a cost — a lower rating, a harder conversation, a smaller bonus, a black mark that resurfaces at layoff time.
- The rational move is to set a target you're already going to hit. Not because the person is dishonest, but because the system is asking them to bet their compensation on their own optimism.
- The organization loses its ability to see reality — the OKR dashboard turns green while the actual trajectory of the business stays flat.
This isn't new. W. Edwards Deming, writing decades before OKRs existed, warned against exactly this pattern in his critique of management-by-objective and numerical quotas: fear-driven target systems, he argued, teach people to manage the appearance of the number instead of the process that produces it. Peter Drucker's original MBO (Management By Objectives) — the direct ancestor of the OKR framework — assumed managers would negotiate honest targets in good faith; it didn't anticipate what happens once HR ties those same targets to compensation review.
The practical upshot: if you want honest targets, look first at what you've coupled to the OKR number, not at how you're training people to write Key Results. If you haven't yet nailed down the difference between a Key Result that tracks a real outcome and one that just tracks activity, our complete guide to advanced OKRs is the place to calibrate that foundation before tackling sandbagging specifically.
Output-Shaped KRs Make Sandbagging Easier to Hide
Sandbagging is far easier to disguise inside an output metric than an outcome metric. "Ship 4 features" can be sandbagged by scoping features down until 4 is trivial. "Increase weekly active usage of the feature by 15 points" is much harder to game quietly, because the number is anchored to something the market controls, not something the team controls.
If your KRs are still written as outputs — launches shipped, features delivered, dashboards built — sandbagging has a hiding place. Our piece on the outcome vs. output OKRs distinction walks through how to rewrite output-shaped Key Results so a safe target becomes visibly safe, instead of quietly safe.
The Tell: Five Signs a Target Was Set to Be Hit
The clearest sign of a sandbagged target is that it sits within a few percentage points of the current baseline — the team is proposing to move a number that's already moving on its own, and calling the drift "ambition." Once you know to look for this, it shows up everywhere in a typical planning deck.
Watch for these five patterns during OKR drafting conversations:
- The target tracks the trend line, not the ambition. If the metric has grown 4% a quarter for the last three quarters and the new KR proposes 5%, that's not a stretch goal — that's extrapolation with a bow on it.
- The KR is phrased as an activity, not a result. "Conduct 10 customer interviews" is unfalsifiable as ambition; it will happen regardless of whether it changes anything. Compare that to "Reduce
time-to-valuefor new customers from 21 days to 10." - The team can't articulate what would make the target hard. Ask "what would have to go wrong for you to miss this?" A confidently sandbagged KR gets a shrug in response — there's no real risk baked in.
- The baseline is fuzzy or missing entirely. When nobody states the current number precisely before proposing the target, that's often not an oversight — an unstated baseline is much harder to challenge.
- Every KR across the team lands in the same 90-100% attainment band, quarter after quarter. Real ambition produces variance: some KRs get blown past, some get missed. Uniform near-perfect scores are a signal of coordinated caution, not coordinated excellence.
| Signal | Sandbagged version | Honest stretch version |
|---|---|---|
| Support resolution time | "Reduce from 14.0 hrs to 13.5 hrs" (already trending down) | "Reduce from 14.0 hrs to 9 hrs, with 11 hrs as an acceptable floor" |
| Feature adoption | "Ship the redesigned onboarding flow" (output, no bar) | "Move day-7 activation from 38% to 55% via the redesigned flow" |
| Sales enablement | "Train 100% of reps on the new pitch" (activity, guaranteed) | "Increase average deal size on trained reps' deals by 20%" |
| Retention | "Maintain churn at or below 4.2%" (current rate, restated) | "Cut logo churn from 4.2% to 3.0% in the SMB segment" |
The table above is worth reading as a pattern, not a checklist: in every sandbagged row, the target restates something already true or already trending. In every honest version, the number requires the team to change what's currently happening, and the team owns a number the market — not their own activity log — will validate.
One more diagnostic worth running: look at how the KR is written relative to the outcome baseline, not just the target. A KR that never states where the metric started is much harder to interrogate than one that opens with "starting from 38%, we'll get to 55%" — the baseline is what makes the ambition (or the lack of it) visible in the first place.
Decoupling OKRs From Judgment: The Psychological Safety Fix
The fix for sandbagging is structural, not motivational — you cannot pep-talk someone into proposing an ambitious target that could cost them their rating. What works is removing the link between the OKR number and any consequence, then rebuilding, deliberately, the trust that missing an honest stretch goal is safe.
Harvard Business School researcher Amy Edmondson, whose work on psychological safety spans three decades of organizational research, defines it as a shared belief that the team is safe for interpersonal risk-taking — including the risk of saying "I'm not sure we'll hit this." Edmondson's research consistently finds that teams without this safety don't actually make fewer mistakes; they just report fewer of them. Sandbagged OKRs are a direct instance of that pattern: not fewer misses, just fewer visible ones.
Three structural moves make decoupling real rather than aspirational:
- Separate the OKR review from the performance review, explicitly and in writing. Google's internal OKR guidance — documented publicly through re:Work and popularized by John Doerr in Measure What Matters — treats OKR attainment as one input among many in a performance conversation, never the formula. State this rule out loud in the kickoff meeting, not just in a policy doc nobody reads.
- Reward honest misses over dishonest hits, visibly. If a team that swung for 60% attainment on a genuinely hard target gets a better outcome in the review cycle than a team that hit 100% on something easy, say so publicly. The first time this happens and people notice, the incentive structure starts to actually shift.
- Kill the goal theater around the quarterly check-in itself. A lot of sandbagging isn't really about the target — it's a defensive reaction to a review ritual that treats the OKR dashboard as a performance trial rather than a planning tool. Our piece on killing goal theater in the quarterly cycle covers how to redesign that ritual so it stops manufacturing the fear that produces sandbagging in the first place.
Cascading Pressure Makes This Worse, Not Better
Sandbagging rarely stays contained to one team. If a VP sandbags a company-level KR to protect themselves, and that number cascades down as a target floor for every team underneath them, the whole org inherits artificially low ambition — and every layer below adds its own margin of safety on top. This is one reason waterfall-style OKR cascades amplify the problem: each level negotiates down from an already-safe number.
Our guide to cascading OKRs without the waterfall trap covers how to link team-level KRs to company outcomes without forcing a top-down number that invites exactly this kind of defensive padding at every layer.
Calibrating Real Stretch: How Ambitious Is Ambitious Enough
A well-calibrated stretch KR should feel uncomfortable to commit to and have a real, non-trivial chance of being missed — somewhere in the range of 60-70% average attainment is the commonly cited healthy band for aspirational goals, not 95-100%. Calibration is a skill, and like any skill it needs a shared reference point the whole team uses the same way.
Doerr's Measure What Matters popularized a two-tier structure that's worth adopting explicitly:
| Goal type | Expected attainment | What a miss means |
|---|---|---|
| Committed goal | ~100% (this is a promise) | Missing signals a real planning or execution failure |
| Aspirational (stretch) goal | ~60-70% on average | Missing is expected and healthy; 100% suggests it wasn't a stretch |
The mistake most teams make is treating every KR as a committed goal, which pushes everyone toward the sandbagging behavior described above. Not every Key Result should be scored the same way — decide, at drafting time, whether a KR is a promise or a swing, and score it accordingly.
Goal-setting research backs the underlying principle even outside the OKR framework specifically. Edwin Locke and Gary Latham's goal-setting theory — built on decades of empirical studies and summarized in their 1990 book A Theory of Goal Setting and Task Performance — found that specific, difficult goals reliably produce higher performance than vague or easy ones, provided the person has genuine buy-in and feedback along the way.
The theory doesn't say "harder is always better" without limit. It says specificity plus real difficulty plus feedback is what drives performance — and that safety plus vagueness is what erodes it.
Anchor Stretch to the Customer, Not the Org Chart
The most durable way to calibrate a stretch target is to anchor it to a customer outcome rather than an internal negotiation. A target that's ambitious relative to last quarter but disconnected from what actually matters to the customer is still theater — it's just harder-looking theater.
Two anchoring questions are worth asking before finalizing any KR:
- What job is the customer hiring this metric improvement to do for them? If you can't connect the number to a real customer job, revisit our Jobs to Be Done complete guide for a sharper way to frame the "why" behind the target.
- Where does this metric sit on the customer's actual experience? Grounding a target in a specific point of friction on the customer journey — rather than an arbitrary percentage bump — makes it much easier to argue for genuine ambition, because the case is "here's a real point of pain we can move," not "here's a number that sounds good."
Before and After: Renegotiating a Sandbagged KR With Shared Risk
Renegotiating a sandbagged target works best as a collaborative recalibration, not an ultimatum — the goal is to move the number up while also giving the owner a credible floor they won't be punished for landing on. Here's a walk-through of what that conversation can look like in practice.
Before (the sandbagged version):
KR: Reduce average customer support ticket resolution time from 14.0 hours to 13.5 hours by end of quarter.
The support lead proposed this after resolution time had already drifted from 14.5 to 14.0 hours over the prior two quarters with no deliberate intervention. The target simply extends a trend that was already happening. Asked what would need to go wrong to miss it, the answer was "nothing, really — we're basically already there."
The renegotiation conversation, structured around three questions:
- "What's the baseline, precisely, and how did we get here?" Surfacing the exact starting number (14.0 hours, drifting down passively) makes the safety of the original target immediately visible to both people in the room — nobody has to accuse anyone of anything.
- "If missing this had zero effect on your rating, what would you actually try for?" This question only works if the decoupling from the previous section is real and the person believes it. The support lead's honest answer, once trust was established: "If we fixed the tier-1 routing logic, I think we could get to 9 hours."
- "What's the floor you'd still call a win, even if the stretch doesn't land?" This is the shared-risk mechanism — it gives the target owner permission to swing hard without betting everything on a single number.
After (the renegotiated version, with shared risk):
KR: Reduce average customer support ticket resolution time from 14.0 hours toward a stretch target of 9 hours, with 11 hours defined in advance as a strong, fully credited result. Manager commits engineering time for the tier-1 routing fix as a dependency.
Notice what changed structurally, not just numerically. The target moved from a 0.5-hour drift to a genuine 3-5 hour improvement. A floor was named explicitly, so landing at 11 hours isn't a miss dressed up as a win — it's a pre-agreed, fully credited outcome. The manager took on a dependency, converting this from "your problem alone" into shared risk, which is exactly the mechanism that makes ambitious commitment safe to offer honestly.
This pattern — real stretch number, pre-named floor, a resourcing commitment from the person requesting the ambition — is the practical shape "decoupling from consequences" takes in an actual drafting conversation. It costs nothing to write down and changes the entire risk calculation for the person proposing the number.
Where Prodinja Fits: Making Sandbagging Visible at Draft Time
Sandbagging is hardest to catch in the moment a KR gets written, because that's exactly when the baseline is freshest in the proposer's head and easiest to quietly omit. A tool that keeps the starting number visible during the drafting conversation removes the option to skip that step.
Prodinja's Outcome baselines are designed to sit next to the target field while a Key Result is being drafted, so the starting number and the proposed number are visible side by side rather than one of them living only in someone's memory. When a target barely moves the baseline — the 14.0-to-13.5-hours pattern from the example above — that gap is easy to see and easy to challenge in the room, before the KR gets locked into a quarterly plan.
It doesn't replace the trust-building work of decoupling OKRs from performance review. It just makes the "is this actually a stretch" question harder to dodge at the exact moment it matters.
Key Takeaways
- Sandbagging is a rational response to a badly designed incentive, not a character flaw — people set safe targets when missing an ambitious one has a personal cost and hitting a safe one doesn't.
- The clearest tell is a target close to current run-rate, phrased as an activity rather than an outcome, with a baseline that's vague or missing entirely.
- Decoupling OKR attainment from performance review is the structural fix; a shared verbal commitment to it only sticks once people see honest misses rewarded over dishonest hits.
- Score committed goals and aspirational goals differently. Roughly 60-70% attainment on a genuine stretch KR is a healthy sign, not a shortfall — 100% across the board is the real warning sign.
- Anchor ambition to a real customer outcome, not an internal negotiation, so the target has a defensible "why" beyond "it sounded stretchy enough."
- Renegotiate with shared risk, not ultimatums: name a real stretch number, pre-agree a credited floor, and have the person requesting ambition take on part of the dependency.
- A visible baseline at drafting time — whether in a spreadsheet or a tool built to surface it automatically — is one of the cheapest ways to make a sandbagged target hard to slip past the room.
Frequently Asked Questions
How do you tell the difference between a sandbagged target and a genuinely conservative one?
Ask what the number was doing before the target was set. A conservative-but-honest target still requires the team to change current behavior to hit it; a sandbagged one just restates the existing trend line. If the metric would land near the target with zero new effort, it's sandbagged regardless of how it's phrased.
Does removing OKRs from performance reviews actually reduce gaming OKR targets?
Directionally, yes — because the core mechanism behind sandbagging is the personal cost of missing, and removing that cost removes the incentive to hide it. This doesn't happen instantly; teams need to see a few cycles where honest misses aren't punished before behavior actually shifts, since trust rebuilds slower than policy changes.
What's a reasonable stretch percentage when setting ambitious OKR targets?
There's no universal number, but the widely referenced benchmark from Google's OKR practice — documented in John Doerr's Measure What Matters — treats roughly 60-70% average attainment on aspirational goals as healthy. Consistently landing at 95-100% across every KR is a stronger signal of overly safe targets than of strong execution.
Should every team be forced to set at least one aspirational OKR?
Mandating a single aspirational KR per team can work as a forcing function, but only if it's explicitly scored differently from committed goals — otherwise it just becomes one more target people quietly sandbag. Pair any mandate with the decoupling and floor-setting practices above, or it will produce the same defensive behavior it's meant to fix.
Is sandbagging always intentional?
Not always — some of it is unconscious anchoring, where a proposer's brain defaults to "what's realistic" instead of "what's ambitious" once they know a rating is on the line, without any deliberate decision to game the system. That's part of why the fix has to be structural (decoupling, visible baselines, shared risk) rather than relying on people simply trying harder to be honest.