Google's 0.6–0.7 average is real, but it only fits one shape of Key Result: a continuous metric scored as a baseline-to-target ratio. Binary launches, guardrail floors, and forward-looking bets need different math. What actually works is choosing the grading system per Key Result type — and locking it in before the quarter starts, not at grading time.
Quick Answer: Google's 0.0–1.0 scale (aim for 0.6–0.7) only fits continuous-metric Key Results. Binary or milestone KRs score a clean 0 or 1, guardrail KRs pass or fail against a floor, and forward-looking bets use a weekly confidence rating instead. Pick the right system per Key Result when you write it, not when you grade it.
Picture the quarter-end review. One Objective, three Key Results — a conversion percentage, a shipped redesign, and a support-ticket ceiling. The first has a clean number to plug into a formula. The other two don't, and the team is about to force all three through the same math anyway.
The Quarter-End Moment Where OKR Grading Falls Apart
OKR grading breaks down at the exact moment a team averages scores from Key Results that were never comparable to begin with — a percentage next to a yes/no launch next to a floor nobody wants to cross. Most OKR training only ever demonstrates the easy case, so nobody notices the mismatch until the numbers refuse to average cleanly.
Take an Objective a lot of product teams would recognize: "Make checkout something customers trust enough to finish." Under it, three Key Results:
- Cart-to-purchase conversion: 61% → 72%
- Ship the redesigned payment form
- Keep payment-related support tickets under 3% of transactions
The first is a textbook Key Result — a number with a baseline and a target. The second is binary: shipped or not. The third is a ceiling, not a summit to climb toward. Run all three through (Actual − Baseline) / (Target − Baseline) and only the first produces an answer that means anything.
Most OKR templates only ever show you the easy case — the one where the ratio just works.
This isn't a rare edge case. A meaningful share of real Key Results are milestones, thresholds, or forecasts wearing a metric's clothing, which means most teams are grading at least one KR per Objective with a formula that was never built for it.
None of this is an argument against Google's formula. It's a reminder that grading is the last step in a process that starts with writing a measurable Key Result in the first place — covered in full in our complete guide to OKRs for product teams.
Why "Just Use Google's 0.6-0.7 Rule" Breaks Down in Practice
Google's ratio formula assumes a Key Result that moves linearly between two numbers you trust — a real baseline, a real target, both measured the same way. The moment a Key Result is a yes/no milestone, a floor you can't cross, or a bet with no historical baseline, the ratio either produces a meaningless number or can't be computed at all.
The formula was never broken. It was just never asked to grade anything but a clean, continuous metric.
The 0.6-0.7 figure itself is real and well-documented. Andy Grove built the underlying intent-based grading approach at Intel, described in High Output Management (1983); John Doerr carried it to Google in 1999 and popularized the specific range in Measure What Matters (2018); Google's own public re:Work guide still codifies it today. None of that history says the ratio applies to every Key Result — only to the ones actually shaped like a metric.
Three situations break the ratio formula specifically:
- Binary or milestone KRs. "Launch the redesigned payment form" has no honest midpoint — it's shipped or it isn't. Forcing a percentage onto it (60% done?) just invents false precision. This is frequently a sign the KR is really an output wearing an outcome's clothing and deserves a rewrite before it deserves a grading formula.
- Guardrail or threshold KRs. "Keep churn under 5%" isn't a target to maximize — crossing the line is a failure regardless of margin, and staying well under it isn't extra credit. It's a pass/fail gate, not a ratio.
- No-baseline discovery KRs. A genuinely new bet often has no historical number to measure from, so any target is a guess dressed up as a decimal until real data exists.
One failure mode is worth naming even briefly: a Key Result that lands at a suspiciously clean 1.0 quarter after quarter. Marissa Mayer, describing Google's own grading culture, is widely quoted saying a perfect score usually means the target was sandbagged, not that the quarter went flawlessly. It's one entry in a much longer list — see our catalog of OKR anti-patterns for the rest.
A Practical Taxonomy: Match the Scoring System to the Key Result Type
Every Key Result falls into one of four scoring shapes: a continuous metric (Google's ratio), a binary milestone (0 or 1), a guardrail threshold (pass/fail), or a forecast bet (a weekly confidence rating, not a retrospective score). Naming the shape before the quarter starts is what makes grading day painless.
| Key Result Type | Scoring System | How It's Calculated | Example | Common Pitfall |
|---|---|---|---|---|
| Continuous metric | Google's 0.0–1.0 ratio | (Actual − Baseline) / (Target − Baseline), clamped 0–1 | Conversion 61% → 72% | Treating a fuzzy proxy metric as if it were precise |
| Binary / milestone | Pass/fail | 1 if shipped or achieved, 0 if not — no partial credit | Ship the redesigned payment form | Inventing a "70% done" score that hides whether it actually shipped |
| Guardrail / threshold | Pass/fail against a floor | 1 if the metric stays inside the floor all quarter, 0 if it crosses even once | Support tickets stay under 3% of transactions | Averaging it into the Objective score instead of treating a breach as a stop-and-review signal |
| Forecast / confidence | Weekly confidence rating | A 0–10 or red/amber/green self-rating, tracked over time, never averaged into the final grade | "How confident are we that we'll hit 500 activated users by quarter-end?" | Confusing a mid-quarter confidence number with the actual end-of-quarter grade |
The first two rows are old news to most OKR programs. The last two are where scoring conversations usually go quiet, because nobody wrote down a rule for them ahead of time.
A Key Result's type isn't a formality — it's the difference between a grade you can defend and a number you made up to fill a cell in a spreadsheet.
Christina Wodtke's Radical Focus (2016) popularized the weekly confidence rating specifically to solve that fourth row — instead of forcing a bet with no baseline into a fake percentage, teams rate their confidence of hitting the target on a simple scale, updated weekly, completely separate from the quarter-end grade.
A confidence rating answers "are we on track." A grade answers "what happened." Neither one is the other.
Where a Key Result sits in the customer journey is often the fastest way to spot which type it is. A retention or expansion metric usually has months of history behind it — a clean baseline. A brand-new acquisition channel or a just-launched feature usually doesn't, which is exactly why it belongs in the forecast row instead of being forced into a ratio it can't support yet.
How to Set Up OKR Scoring Before the Quarter Starts
Setting up scoring correctly is front-loaded work, done in the same meeting where the Key Result gets written — not a scramble at quarter-end. Five moves, done in order, mean grading day is a five-minute read of numbers everyone already agreed on, not a debate about what the number even means.
- Classify the type before you argue about the target. Ask whether the Key Result is a continuous metric, a milestone, a guardrail, or a forecast — the type determines the formula, and arguing over a target number before agreeing on the type wastes the conversation.
- Write the formula into the OKR document itself. Next to the Key Result, name the exact scoring rule — not as tribal knowledge one person remembers, but as a visible line anyone reviewing the doc in twelve weeks can apply without asking around.
- Anchor the metric to a real customer job, not an internal proxy. A Key Result scored against a number nobody outside the team recognizes is easy to game quietly. Grounding it in an actual job the customer is hiring the product to do — the same discipline behind Jobs-to-Be-Done — keeps the score honest.
- Separate the weekly confidence check-in from the quarter-end grade. One is forward-looking and informal; the other is backward-looking and final. Reporting both on the same 0.0–1.0 scale is how a merely rough week turns into a false failing grade.
- Calibrate scores as a group, not solo. A five-minute round where each Key Result owner explains their number out loud catches quiet grade inflation before it becomes a habit — the same reason code review exists.
Weighting Key Results and Reading the Final Score Honestly
Not every Key Result under an Objective deserves equal weight in the average — a guardrail that only ever fails matters differently than a stretch metric that's 80% of the way there. Decide the weighting the same day you decide the scoring type, and read the final number in context, not as a single verdict on the quarter.
| Dimension | Weekly Confidence Check-in | Quarter-End Grade |
|---|---|---|
| Purpose | Forecast — flag risk while there's time to react | Record — settle what actually happened |
| Timing | Every week or two, throughout the cycle | Once, at cycle close |
| Typical scale | 0–10 or red/amber/green, self-reported | 0.0–1.0, calculated or calibrated |
| Who sets it | The Key Result owner, informally | The team, ideally calibrated together |
| What a low number means | Time to escalate or adjust course | Data for next quarter's targets, not a punishment |
Weighting matters because a simple average can hide the one Key Result that actually mattered. Most dedicated OKR platforms — Betterworks, Perdoo, and Quantive among them — let teams assign uneven weights to Key Results under one Objective rather than averaging them straight, precisely because a 30%-weighted guardrail and a 70%-weighted growth metric shouldn't count the same.
A quick illustration, scoring three Key Results at 0.9, 0.6, and 1.0:
- Flat average: (0.9 + 0.6 + 1.0) ÷ 3 = 0.83
- Weighted 20/50/30, because the middle KR is the actual stretch goal the Objective depends on: (0.9×0.2) + (0.6×0.5) + (1.0×0.3) = 0.78
The gap looks small, but it's the difference between a report that flatters the quarter and one that reflects what the Objective actually depended on. This isn't just OKR-specific opinion, either — decades of goal-setting research from psychologists Edwin Locke and Gary Latham associate specific, challenging targets paired with regular feedback with meaningfully better performance than vague goals. A score is the feedback half of that equation, and it only works if the number means something specific.
One more caution: don't compare raw scores across teams without asking how each team's Key Results actually cascaded from the company goal — a platform team's guardrail-heavy scorecard will look more conservative than a growth team's stretch-heavy one by design, a difference in how OKRs cascade, not in effort or execution quality.
The Prodinja Angle: Grading Follows the Same Discipline as Prioritization
Ranking a feature in that workspace means naming three things before it can be scored at all:
- Reach — how many users or accounts the change actually touches
- Impact — the specific metric it's meant to move, and roughly by how much
- Confidence — how sure that estimate really is, not a guess dressed as a number
Key Takeaways
- Google's 0.6-0.7 average only fits one Key Result type — a continuous metric with a real baseline and target. It was never meant to grade a milestone, a guardrail, or a no-baseline bet.
- Classify each Key Result's type before the quarter starts: continuous metric, binary milestone, guardrail threshold, or forecast confidence — the type determines the formula, not the other way around.
- Binary and guardrail KRs need pass/fail scoring, not an invented percentage — false precision on a yes/no outcome hides more than it reveals.
- Weekly confidence ratings and quarter-end grades are two different systems serving two different purposes; conflating them turns a rough week into a false failing grade.
- Weight Key Results unevenly when they don't deserve equal say in the Objective's average — most dedicated OKR platforms support this for a reason.
- Calibrate scores as a group, not solo — a short shared review catches quiet grade inflation before it becomes the norm.
- A scoring mechanism is what separates a Key Result from a wish — the same discipline that makes a prioritized backlog item defensible instead of a guess.
Frequently Asked Questions
What is a good OKR score?
A good OKR score depends on the Key Result's type: a continuous metric scores well around 0.6-0.7, a binary milestone is simply hit (1) or missed (0), and a guardrail is a clean pass or fail against its floor. There's no single "good number" that applies across all three at once.
How do you score a Key Result that isn't a number, like a launch or milestone?
Score it as pass/fail — 1 if it shipped or was achieved, 0 if it wasn't — instead of inventing a percentage like "70% done." A milestone Key Result has no honest midpoint, and forcing a ratio onto it usually means the KR is really an output that should be rewritten as a measurable outcome first.
Is a weekly confidence rating the same as an OKR grade?
No — a confidence rating is a forward-looking, informal forecast updated throughout the quarter, while an OKR grade is a backward-looking, final number set once at cycle-close. Popularized by Christina Wodtke's Radical Focus, the confidence rating exists to flag risk early; the grade exists to record what actually happened.
Should every Key Result under one Objective be weighted equally?
Not necessarily — a guardrail Key Result that only exists to prevent regression often deserves less weight in the average than a stretch metric the Objective actually depends on. Decide the weighting when you write the Key Results, using the same judgment you'd apply to any prioritization call, not after the numbers come in.
How often should OKRs be graded?
Grade formally once, at the end of the cycle — but check in on confidence weekly or biweekly throughout it. Grading more often than that turns a goal-setting tool into a status-report ritual, while grading less often removes the chance to course-correct while it still matters.