Annual planning turns into theater when it's structured as a negotiation over promises rather than an allocation of finite capacity. The fix is to plan in outcome bets sized against real team-capacity, reconcile top-down ambition with bottom-up reality, and rank a portfolio instead of stacking a wish list. What survives that process becomes a living spec, not a slide.
Quick Answer: Run annual planning as capacity allocation, not promise-making. Set a small number of outcome bets, size each against actual team-capacity (not calendar dates), reconcile leadership's top-down ambition with teams' bottom-up reality, and rank what's left the way an investor sizes a portfolio.
Why Annual Planning Becomes Roadmap Theater
Annual planning becomes theater when stakeholders submit feature wish lists, product teams pad estimates to survive the haggling, and the resulting document is a set of dated promises dressed up as strategy. Naming this failure mode is the first step — most VPs recognize it the moment it's described, because they've run it, or sat through it, more than once.
The ritual is familiar. Sales asks for the three features closing this quarter's biggest deals. Support forwards a backlog of "quick fixes" that never turn out to be quick. Finance wants a cost line for every initiative before it's scoped. Each function optimizes for its own outcome, and the roadmap that emerges is really a truce — not a strategy.
Three structural problems make this worse than a scheduling headache:
- Dates get treated as commitments even though nobody has validated the underlying assumption yet — turning a hypothesis into a deadline before anyone's tested it.
- Feature line items hide the actual outcome a stakeholder is chasing, so two requests that serve the same job get built twice, or neither gets built well.
- Estimation becomes political because teams learn that padded numbers survive negotiation and honest numbers get cut, which quietly rewards the wrong behavior year after year.
There's a real, if uncomfortable, data point behind why so many of these promises turn out not to matter. The Standish Group's long-running CHAOS research on software features has repeatedly found that a large share of shipped functionality — historically estimated somewhere in the 40-60% range depending on the survey year — goes rarely or never used by customers. If close to half of what a wish-list roadmap promises won't move outcomes anyway, the negotiation was never the right unit of planning. This is one reason the VP of Product role carries such different pressure than the director role beneath it — the accountability for what the org actually builds sits with the VP, as covered in our complete guide to the VP of Product role.
Reframe the Exercise: Outcomes and Bet Sizes, Not Dated Features
Annual planning should answer "what outcomes are we betting on, and how big a bet are we willing to make" instead of "what will we ship and by when." Replace feature line items with a small portfolio of outcome bets, each sized in team-capacity rather than a calendar date, so the plan survives contact with reality.
This isn't a novel idea — it's the same discipline Basecamp's Ryan Singer formalized in Shape Up under the name appetite: instead of asking "how long will this take," you ask "how much time is this worth to us," fix that budget, and shape the scope to fit inside it. Applied to annual planning, appetite becomes the unit of the whole exercise. You're not committing to ship the permissions rebuild by March; you're committing at most one team-quarter to it, and if it doesn't fit, the scope shrinks, not the deadline.
Bet sizes replace date commitments
A five-tier bet scale gives leadership and teams a shared vocabulary for "how big a swing is this," independent of which quarter it lands in:
| Bet Size | Capacity Commitment | Typical Scope | Example |
|---|---|---|---|
| XS | Under 0.25 team-quarter | Single team, quick win, low risk | Fix the top drop-off step in onboarding |
| S | ~0.5 team-quarter | Single team, partial quarter | Ship a self-serve upgrade flow |
| M | 1 team-quarter | Single team, full quarter | Rebuild the permissions model |
| L | 2-3 team-quarters | Multi-team, half a year | New billing platform |
| XL | 4+ team-quarters | Multi-team, full year or more | Platform migration or new product line |
The table's real value is the conversation it enables: a stakeholder who wants an "L" bet delivered in an "S" window has just revealed that the ask isn't actually sized yet — the disagreement is now about scope, not about whether the team is slow.
Plan in outcomes, track in outputs
Each bet should carry an outcome statement — the metric or customer behavior it's meant to move — separate from the output that will actually ship. Reduce time-to-first-value in onboarding is the outcome; the self-serve upgrade flow is one output that might move it. Keeping the two separate is what lets a team change the output mid-quarter without breaking the annual commitment, because the commitment was never to the feature.
Run the Top-Down/Bottom-Up Reconciliation
Reconciliation works by running two independent estimates — leadership's top-down strategic envelope and teams' bottom-up capacity reality — and then closing the gap between them in the open, rather than letting one side silently overrule the other. Skipping this step is what produces a roadmap that looks agreed-upon in the room and falls apart in week three.
Step 1: Leadership sets the top-down envelope
Start with an allocation across horizons, not a list of projects. McKinsey's long-standing Three Horizons framework — allocating investment across the core business (Horizon 1), emerging opportunities (Horizon 2), and speculative future bets (Horizon 3) — is a useful top-down scaffold precisely because it forces a percentage conversation before a project conversation. Most orgs over-index Horizon 1 by default; naming a target split (even a rough 70/20/10) makes that bias visible and debatable instead of accidental.
Step 2: Teams report bottom-up flow distribution
Ask every team to report, honestly, how its capacity actually splits today — not how leadership assumes it splits. Mik Kersten's Flow Framework (from Project to Product) gives a clean vocabulary for this: capacity flows into features, defects, risk (security, compliance, tech debt with a ticking clock), and debt (everything else that slows future work). Teams that have never measured this typically discover far less capacity is free for new bets than leadership assumed — often only 50-70% once keep-the-lights-on work, on-call, and ramp time for new hires are subtracted honestly.
Step 3: Reconcile the delta in the open
The gap between the top-down envelope and the bottom-up reality is the actual planning conversation — everything before this was just data-gathering. Put both numbers in the same room and negotiate the delta visibly: cut scope, add headcount, extend timeline, or explicitly accept more Horizon-1 debt for one more year. This is the same alignment problem behind getting multiple executives to commit to a single shared roadmap instead of each defending their own version, which is why we've written separately about driving exec alignment around one roadmap. Bain's RAPID framework (Recommend, Agree, Perform, Input, Decide) is worth borrowing here too, if only to name who actually holds the decision when the top-down and bottom-up numbers don't match — usually you, as VP, with input from Engineering and Design leadership rather than a vote.
Hold the Capacity Truth Conversation With Peers
A capacity truth conversation is a peer-level meeting — VP Product, VP Engineering, and design/research leadership — that surfaces the real, unpadded capacity available for new bets before any prioritization begins, rather than after a roadmap has already been promised externally. Held too late, it becomes a fight over what to cut from a plan people have already told the board about.
Who's in the room and what they owe each other
Three commitments make this conversation work instead of degenerating into another negotiation:
- Engineering brings the real flow-distribution number — the percentage actually available for new bets after debt, defects, and on-call, not the number that makes the roadmap fit.
- Product brings the ranked bet portfolio, not a wish list, so the conversation is "which bets fit" rather than "everything, somehow."
- Design and research bring discovery capacity constraints too — a bet with no design or research bandwidth behind it isn't actually funded, whatever the engineering allocation says.
Translating the gap for the board
When the honest capacity number is smaller than what leadership has implicitly promised upward, that gap needs a business narrative, not a confession. Framing a reduced Horizon-1 investment as a deliberate trade for Horizon-2 growth, backed by the bet-sizing table, reads as strategy; framing it as "we're behind" reads as failure. That reframing is exactly the skill covered in our guide to building the board-ready product business narrative — the capacity truth conversation is where you build the material for it, before you're standing in front of the board improvising.
Teresa Torres, in her work on continuous discovery, makes a related point: teams that treat discovery capacity as a rounding error end up shipping the wrong bet confidently. The capacity truth conversation should protect discovery time the same way it protects engineering time.
Worked Example: From Wish List to Ranked Bet Portfolio
Converting a stakeholder wish list into a ranked bet portfolio means running every request through three filters — the job behind it, a defensible score, and a capacity-fit bet size — then ranking what's left against the envelope from the reconciliation above. The example below shows the same ten-item wish list before and after.
The raw wish list
A typical Q4 planning inbox looks something like this, verbatim: "add SSO," "fix the slow dashboard," "build a mobile app," "let enterprise customers white-label," "add bulk export," "redesign onboarding," "add Slack integration," "improve search," "add usage-based billing," "support SCIM provisioning." Ten asks, no shared scale, no visible connection to strategy.
Step 1: Find the job behind each request
Before scoring anything, ask what job the stakeholder is actually hiring the feature to do — the same discipline covered in our complete guide to Jobs to Be Done. "Add a mobile app" often decomposes into a narrower job like "let a field rep check status without opening a laptop," which might be solvable with a lighter-weight response than a native app. Naming the job before scoring prevents the portfolio from over-funding the loudest request instead of the most valuable one.
Step 2: Score with reach, impact, confidence, effort
Impact scoring gets easier once you can point to exactly where in the customer experience the friction sits — which is why mapping the journey and its emotion curve first, as we cover in our complete guide to customer journey mapping, makes the scoring conversation shorter, not longer. The table below applies a RICE-style pass (Reach, Impact, Confidence, Effort) to the same ten requests.
| Request | Job Behind It | RICE Score (relative) | Bet Size | Decision |
|---|---|---|---|---|
| Fix the slow dashboard | Trust the data is current before a client call | High | XS | Fund now |
| Redesign onboarding | Get value before the trial expires | High | M | Fund now |
| Add SSO | Pass procurement security review | High | S | Fund now |
| Add usage-based billing | Match spend to actual value received | Medium-High | L | Fund, next window |
| Improve search | Find existing work without asking support | Medium | S | Fund now |
| Support SCIM provisioning | Automate offboarding for large accounts | Medium | S | Defer to next cycle |
| Add bulk export | Get data out for one internal report | Low-Medium | XS | Defer |
| Add Slack integration | Get notified without checking the app | Low-Medium | S | Defer |
| Let enterprise customers white-label | Resell under their own brand | Low (narrow reach) | L | Kill for now |
| Build a mobile app | Check status away from a desk | Low (job solvable smaller) | XL | Kill, revisit narrower job |
Reading the table left to right shows the shift in logic: reach and job clarity, not who asked loudest, decided what got funded. The two "kill" rows aren't rejections of the underlying need — they're a bet-sizing verdict that the job can likely be served by something smaller than what was literally requested.
Step 3: Rank against the capacity envelope
Lay the funded rows against the top-down/bottom-up envelope from the reconciliation step and stop funding once the envelope is full — not once the list runs out. If the envelope only covers four of the five "fund now" bets, that's the real capacity truth conversation happening in miniature, on paper, before it happens with peers in a room.
Where a Scoring Tool Helps (and Where It Doesn't)
Scoring isn't the hard part of this process — the hard conversations about capacity and trade-offs are. But a consistent, visible scoring basis makes those conversations shorter, because nobody's arguing about whether the method itself is fair.
This is the part of the process Prodinja is built to support without pretending to replace judgment. Its RICE and Kano prioritization gives every wish-list item the same defensible scoring basis shown in the table above, so a ranked bet portfolio isn't just one VP's gut call dressed up in a spreadsheet. And once bets are funded, Spec Studio turns the surviving outcomes into a living PRD — with PR-style diffs and readiness gates — instead of letting the decision die in a planning slide deck that nobody opens again after the offsite.
Key Takeaways
- Annual planning fails as a wish-list negotiation because dated feature promises get treated as commitments before anyone's validated the underlying assumption.
- Plan in outcome bets sized by capacity, using a scale like XS-XL team-quarters, so scope — not the deadline — is what flexes when reality intrudes.
- Reconcile top-down strategy with bottom-up flow distribution explicitly, using frameworks like McKinsey's Three Horizons and Kersten's
Flow Framework, rather than letting one side silently win. - Hold the capacity truth conversation before promising anything externally, with Engineering and Design bringing their real, unpadded numbers to the table.
- Convert wish lists into ranked portfolios by finding the job behind each request, scoring it consistently, and sizing it as a bet before ranking against the actual envelope.
- A defensible scoring basis shortens the hard conversations about trade-offs — it doesn't replace them.
Frequently Asked Questions
How is capacity-based roadmap planning different from a normal roadmap?
Capacity-based planning allocates a portfolio of outcome bets against real, measured team-capacity, while a normal roadmap lists dated features against an assumed, usually inflated, capacity. The difference shows up the first time reality intrudes: a capacity-based plan flexes scope, while a feature roadmap breaks its dates.
How many bets should a product org fund in an annual plan?
Most orgs get more value from funding fewer, bigger, well-sized bets than from spreading capacity across a long list of small asks. A useful check is whether every "fund now" bet fits inside the reconciled capacity envelope from the top-down/bottom-up process — if it doesn't fit, cut a bet rather than shrinking every bet a little.
What's the right cadence for revisiting an annual plan?
An annual plan should set the outcome bets and capacity envelope, then get revisited at a quarterly cadence to check bet sizes against actual flow distribution. Waiting a full year to check assumptions defeats the purpose of sizing in capacity rather than dates in the first place.
How do you say no to a stakeholder's feature request without damaging the relationship?
Show the job behind their request, the score it received relative to other funded bets, and the capacity envelope that decided the cutoff — the "no" then reads as a portfolio decision, not a personal rejection. This is much harder to do credibly without a consistent scoring basis stakeholders have already seen applied to other requests.
Does capacity-based planning still work for a company scaling past its founder-led planning era?
Yes, and it becomes more necessary, since informal, tribal-knowledge planning breaks down exactly when an org gets too large for one person to hold the whole picture in their head. That transition — building a planning process that survives without you personally chasing every team — is the operating-model problem we cover in building a product operating model that runs without you.