You're actually right more often than you think only if your confidence tracks your hit rate — most PMs aren't calibrated, they're either chronically overconfident (a "70%" that lands 45% of the time) or falsely humble. The fix isn't better predictions; it's a system that scores every forecast against its stated confidence, monthly, so calibration becomes visible instead of assumed.
Quick Answer: Calibration means your stated confidence matches your actual hit rate — a "70% confident" call should come true roughly 7 times in 10. Tag every prediction with a confidence level, review outcomes on a fixed monthly cadence, and compute a simple
Brier scoreto see whether you're overconfident, underconfident, or genuinely well-calibrated.
What Calibration Actually Means (and Why Most PMs Aren't)
Calibration is a measurable property, not a personality trait: across every time you said "70% confident," about 70% of those predictions should have actually come true. Most people's gut-feel confidence is systematically off — usually overconfident on hard, novel calls and underconfident on easy, familiar ones — because intuition has no built-in feedback loop.
The idea comes from forecasting research, most visibly Philip Tetlock and Dan Gardner's Superforecasting (2015), which grew out of the multi-year, IARPA-funded Good Judgment Project. Ordinary volunteers who tracked and scored their own predictions became directionally more accurate over time than professional intelligence analysts working with classified information — not because they were smarter, but because they measured themselves against reality instead of trusting their gut.
Product management runs on exactly this kind of forecast, just dressed up as a roadmap decision:
- "This redesign will lift activation by 10-15%."
- "Feature adoption will clear 30% of eligible users in 60 days."
- "Churn won't move if we cut the free tier's storage limit."
- "This one-way-door pricing change is worth the risk."
Every one of these is a prediction with an implicit confidence level. Most PMs never write the confidence down, and even fewer check back. That's the gap this whole practice closes.
This gap isn't unique to product management. Doug Hubbard's calibration-training research, described in How to Measure Anything, found that most professionals' "90% confidence" estimates only turn out correct roughly 50-70% of the time — and that a few hours of structured practice with feedback measurably narrows that gap. PMs face the same fixable problem, just wearing a roadmap instead of a spreadsheet.
Why the Gap Persists
Three forces keep PMs uncalibrated even when they're smart and experienced:
- No scoring loop. Roadmap bets resolve months later, after attention has moved to the next quarter.
- Hindsight bias. Once you know the outcome, memory quietly edits what you "really" thought going in.
- No unit of measurement. Without a number attached, "I was pretty confident" can mean anything from 55% to 95%.
The Simple Practice: Tag, Predict, Review
The practice is three repeatable steps: state a specific, falsifiable prediction with a metric and a threshold; attach a confidence percentage at the moment you decide, not afterward; and revisit the prediction on a fixed date to score it right, partly right, or wrong. Doing this monthly turns gut calls into a dataset.
Here's the loop in practice:
- Write the prediction as a testable claim. Not "this will help retention" but "
day-30 retentionfor the new cohort will exceed 42%." Vague predictions can't be scored, which is exactly why they feel safer to make. - Attach a confidence number at decision time. Pick from 50%, 60%, 70%, 80%, or 90% — five buckets are enough resolution without inviting false precision.
- Set a review date when you log it, not when it's convenient later. A 60-day feature bet needs a day-60 review date stamped on day zero.
- Score the outcome on the review date: right, partly right, or wrong against the threshold you wrote down, not against a softened version of it.
- Batch review monthly. Pull every prediction that came due in the last 30 days and score it in one sitting — a habit that pairs naturally with the kind of recurring reflection covered in the complete guide to PM craft reflection.
A Worked Example
Say you're deciding whether to collapse a three-step signup flow into one step. You predict: "Signup completion rate exceeds 55% within 30 days of shipping." You log 75% confidence at decision time and set a review date for day 31.
On day 31, completion lands at 52%. That's a partly right call — directionally correct, just short of the threshold you actually wrote down. Score it honestly instead of rounding up to "basically right," or the whole log quietly loses its value.
What Makes a Prediction Worth Tracking
Not every hunch deserves a slot in the log. The best candidates are predictions tied to a decision you're actually making, with a number attached and a date you'll hit regardless of what else is on fire that week.
Good sources of trackable predictions include:
- Feature adoption forecasts — will the feature clear a specific usage threshold, especially one framed around a customer job as covered in the complete guide to jobs-to-be-done?
- Metric-movement bets — will a specific funnel metric move by X% after a change ships?
- Journey-stage predictions — will drop-off at a particular stage of the customer journey shrink after a fix?
- Reversibility calls — will a one-way-door decision hold up in six months?
A Starter Calibration Table
A calibration table groups your predictions by confidence bucket and compares stated confidence against actual hit rate — the clearest way to see whether "70%" means what you think it means. If the two columns roughly match, you're calibrated; a persistent gap in one direction reveals a specific bias to correct.
Here's a starter table you can copy into a spreadsheet or journal, seeded with the shape a healthy first few months of logging might take:
| Confidence you stated | Predictions made | Actual hit rate | Gap |
|---|---|---|---|
| 50-59% | 8 | 50% | On target |
| 60-69% | 12 | 58% | Slightly overconfident |
| 70-79% | 15 | 60% | Overconfident |
| 80-89% | 10 | 70% | Overconfident |
| 90-99% | 6 | 83% | Overconfident |
This particular pattern — a widening gap as stated confidence climbs — is one of the most common calibration failures among forecasters of all kinds, not just PMs. It shows up so consistently in calibration research, echoed in Daniel Kahneman's writing on professional overconfidence in Thinking, Fast and Slow, that "assume you're overconfident at the top end" is a reasonable default until your own table says otherwise.
Reading Your Own Table
Once you have 20-30 resolved predictions, look for the shape, not any single row:
- Flat and matching across buckets means you're genuinely calibrated — rare, and worth noticing when it happens.
- Gap grows at high confidence means you're overconfident on your strongest convictions — the costliest pattern, since those are the bets you'd stake the most on.
- Gap grows at low confidence (your "50%" calls hit 70% of the time) means you're underselling calls you should trust more.
Setting Up Your Own Table
Five columns are enough to start: the prediction, the confidence you logged, the review date, the outcome, and the resulting score. A basic spreadsheet works from day one — sort by confidence bucket once you have enough rows for the grouping to mean anything.
Resist the urge to add more columns before you have data. Category, stakes, and reversibility are useful tags once the practice is established, but an empty table with fifteen columns is a good way to never actually start.
Brier Score, Kept Simple
A Brier score is a single number that measures how far your confidence sat from the actual outcome, averaged across predictions — 0 is perfect, 0.25 is what pure coin-flip guessing produces, and 1 is confidently, completely wrong. You don't need the full formula memorized; you need the version below.
The scoring rule, named for meteorologist Glenn W. Brier who introduced it in 1950 for weather forecasts, is:
Brier score = (confidence as a decimal − outcome)²
Where outcome is 1 if the prediction came true and 0 if it didn't. Average the scores across all your predictions for a running total.
| Prediction | Confidence stated | Outcome | Brier score |
|---|---|---|---|
| Adoption clears 30% | 80% (0.8) | True (1) | 0.04 |
| Churn holds flat | 60% (0.6) | False (0) | 0.36 |
| Retention beats 42% | 90% (0.9) | True (1) | 0.01 |
| Onboarding lift >10% | 70% (0.7) | False (0) | 0.49 |
Average of these four: 0.225 — close to coin-flip territory, driven almost entirely by the two high-confidence misses. That's the real value of the score: it punishes confident wrong calls far more than confident right ones, which is exactly the asymmetry that should worry a PM staking a roadmap on conviction.
A
Brier scoreunder roughly 0.20 across a few dozen predictions is a reasonable "you're doing fine" line for an individual PM tracking their own roadmap bets; above 0.30 usually means either your predictions are too vague or your confidence is inflated.
Scoring a "Partly Right" Outcome
The classic Brier score assumes a binary outcome, but real product bets rarely resolve that cleanly. A practical workaround: score "partly right" outcomes as 0.5 instead of 1 or 0 in the formula — it keeps the math honest without pretending your world is more binary than it is.
This is a deliberate simplification, not the multi-outcome version professional forecasters use. For a personal practice, simplicity you'll actually maintain beats precision you'll abandon after two months.
Confidence Isn't the Same as Certainty
A well-calibrated 60% is genuinely useful — it just means you should be right about 6 times in 10, not that you're hedging. The goal isn't to inflate every number toward 90%; an overconfident 90% is worse than an honest 60%, because it's wrong more often than the label admits.
Some of your best-calibrated predictions will sit at 55-65% forever, because the underlying decision is genuinely uncertain. That's fine. The Brier score doesn't reward bravado — it rewards accuracy at whatever confidence level you actually hold.
Where Calibration Breaks: Common Traps
Calibration data is easy to collect and easy to misread — the traps aren't in the math, they're in the habits around it. Knowing the four most common failure modes ahead of time prevents months of quietly meaningless data.
- Hindsight-adjusted logging. Going back to "clarify" a prediction after you already know the outcome silently destroys the exercise — this is precisely the failure mode addressed in the case for decision logs over hindsight bias. Log the confidence once, at decision time, and never touch it again.
- Rounding to comfortable numbers. Everyone's tempted to write "50%" or "80%" because they're round. Force yourself to consider the full range; a habit of only ever writing 70% just means you've replaced a hunch with a fake-looking number.
- Small-sample overconfidence in the table itself. Five resolved predictions in the 90% bucket tell you almost nothing — treat any bucket under roughly 15 predictions as provisional, not proof.
- Cherry-picking which predictions get logged. Skipping the log on days a bet feels shaky quietly removes your worst misses from the data before you can learn from them.
Vague Predictions Are the Root Cause of Most of These
Nearly every trap above traces back to one root problem: a prediction that wasn't specific enough to score cleanly in the first place. "This will probably help" isn't a prediction — "day-30 retention exceeds 42%" is. The discipline of writing a scorable claim is most of the work; the confidence number and the review date are just bookkeeping once that's done.
None of these traps take willpower to avoid — they take structure. A fixed template of prediction, confidence, and review date removes the decision of "should I round this?" or "should I skip logging this one?" before it can happen.
Make It a Habit, Not a One-Off Audit
Calibration only becomes useful as a repeated practice, not a single retrospective exercise — the value comes from watching your own gap narrow, or a specific bias persist, across dozens of predictions over months. That requires the same kind of recurring cadence covered in deliberate practice for PMs. A monthly 20-minute review is enough to keep the dataset alive.
Pair it with a lighter-weight running log of the frictions and assumptions that prompted each prediction in the first place — the kind of habit described in building a friction-journal habit for product sense — so you can trace why a call was overconfident, not just that it was.
Calibration Scales to a Team, Too
The same practice works across a product team, not just one PM — compare calibration tables side by side and patterns emerge fast: who's systematically bold, who's systematically cautious, and whose 80% you can actually take to the bank. That context is often more useful in planning than any single prediction on its own.
Team-level calibration also tends to surface where confidence and seniority diverge. A recurring finding across structured-forecasting research, including Tetlock's tournaments, is that experience predicts domain knowledge far more reliably than it predicts calibration.
Where Prodinja Fits
Calibration doesn't need a perfect system to start paying off — it needs one tagged prediction and a date to look back at it. The table, the score, and the pattern-spotting only exist because that first entry did.
Key Takeaways
- Calibration means confidence matches hit rate — a "70% confident" call should come true about 7 times in 10, not more, not less.
- Write predictions as testable claims with a specific metric, threshold, and review date, or they can't be scored at all.
- Attach confidence at decision time, never after the fact — retroactive confidence is hindsight bias wearing a number.
- A calibration table by confidence bucket reveals whether you're overconfident, underconfident, or genuinely well-calibrated, and where.
- A simple
Brier score—(confidence − outcome)², averaged — turns "how did I do" into one trackable number; under roughly 0.20 is a reasonable target. - Review monthly, not sporadically — the pattern only shows up once you have 20-30 resolved predictions to look at.
- The most common failure is overconfidence at the high end — treat your strongest convictions with the most skepticism, not the least.
Frequently Asked Questions
What is a good Brier score for a product manager to aim for?
A Brier score under roughly 0.20 across a few dozen predictions is a reasonable target for an individual PM tracking roadmap bets; 0.25 is what pure random guessing produces, so anything near that suggests your predictions aren't adding information. Scores climb fastest when a handful of high-confidence calls turn out wrong, so those deserve the closest post-mortem attention.
How many predictions do I need before I trust my calibration?
You need roughly 20-30 resolved predictions per confidence bucket before the hit rate in that bucket is meaningful, and closer to that number overall before trusting the shape of your full table. Fewer than that and a single unlucky or lucky call can swing the percentage dramatically — treat early data as provisional, not diagnostic.
Is calibration the same thing as being a good forecaster?
No — calibration and accuracy are related but distinct: a forecaster can be perfectly calibrated while still being unhelpful if every prediction sits at 50%. The goal is calibrated confidence at useful extremes, meaning you're willing to say 85% or 15% when you actually know something, and your hit rate backs that number up.
Do I need to track every single decision I make?
No — track decisions that are specific, falsifiable, and worth the discipline, typically bigger roadmap bets, pricing or positioning calls, and one-way-door decisions rather than routine day-to-day choices. A handful of well-chosen predictions a month, tracked consistently, teaches you more than dozens logged sporadically and never reviewed.
What's the difference between a calibration table and a Brier score?
A calibration table shows the pattern across confidence buckets — where your overconfidence or underconfidence concentrates — while a Brier score compresses everything into one number for tracking trend over time. Use the table to diagnose where you're miscalibrated, and the score to see whether you're improving overall.