A good North Star metric passes three tests: it reflects real customer value, it moves before revenue does (leading, not lagging), and it resists being gamed by short-term tricks. Pick one that fails any test and you'll optimize your team into incentivizing the wrong behavior — often without noticing for months.
Quick answer: Test a candidate metric against three filters — customer value, leading indicator, gaming resistance — before you anchor a team's roadmap to it. A metric that passes only one or two will quietly reward the wrong work.
What a North Star Metric Actually Does (And Why Most Picks Fail)
A North Star metric is not a KPI you glance at in a dashboard — it's the single number your team organizes roadmap decisions, experiment prioritization, and trade-off calls around. Most picks fail because teams choose a metric that's easy to measure, not one that's honest about value. The failure shows up slowly, as a gap between "the number went up" and "customers are better off."
The mental shift that separates a durable North Star from a vanity one is simple to state and hard to internalize: a metric is an incentive system, not just a number. The moment you tell a team "grow this," every person downstream starts finding the cheapest path to move it — including paths that have nothing to do with the value the metric was supposed to represent. Choosing a North Star metric is really choosing what your organization will optimize for on autopilot.
This is also why North Star selection sits squarely in a growth PM's job description rather than being a side project for analytics. As the growth PM role complete guide lays out, growth PMs are accountable for the metric architecture underneath a product, not just the experiments run against it — and the two responsibilities differ from the artifact-driven habits covered in the growth PM vs. core PM skill stack comparison. A core PM ships a roadmap; a growth PM has to defend a number's integrity in front of a whole company.
Three warning signs suggest a team is drifting toward a vanity North Star:
- The metric can go up while a plausible customer complaint about the product also goes up.
- Nobody on the team can explain, in one sentence, why the metric causes revenue or retention rather than merely correlating with it.
- The metric was chosen because a competitor reports it publicly, not because it fits this product's value delivery.
The Three-Part Test: Value, Leading, Hard to Game
A candidate North Star metric earns the role only if it clears three independent filters at once — reflects real customer value, moves before your lagging business outcomes do, and resists being inflated without real value being created. Passing two out of three isn't a pass; a metric that's directional but gameable will get gamed eventually.
Reflects Real Customer Value
The metric should only be able to increase when a customer genuinely got more value from the product. This is where a Jobs to Be Done lens is more useful than a funnel diagram: value is defined by the job the customer hired your product to do, not by the interaction your interface happens to log. The Jobs to Be Done complete guide is a useful gut-check here — if you can't map your candidate metric to progress on a customer's job, it's measuring activity, not value.
A quick test: describe the metric to a customer in plain language and ask if they'd agree that a higher number means they're better off. If a customer would shrug — or worse, wince — the metric is measuring the business's convenience, not theirs.
Leading, Not Lagging
Revenue, churn, and NPS are lagging — they tell you what already happened. A North Star metric needs to move earlier in the causal chain, ideally at the point in the customer journey where value first becomes real. That's what makes it actionable: a team can push a leading number this quarter and expect the lagging ones to follow, rather than waiting on a metric that only confirms decisions made months ago.
Mapping a candidate metric against the customer journey complete guide framework is a fast way to check this. If your candidate sits at the very end of the journey — a renewal, a annual contract value — it's a business outcome, not a North Star; a real North Star usually sits at the "aha" or habit-formation stage, closer to where value is first delivered.
Hard to Game
The hardest of the three to verify upfront, because gaming a metric usually looks identical to genuinely improving it — right up until it doesn't. DAU climbs whether people are getting value or being manipulated by a notification loop; signups climb whether the funnel got easier or the qualification bar got lowered. The test isn't "can this number be moved," every number can be moved — it's "can this number be moved without the underlying value also moving."
| Test | What it checks | Fails when |
|---|---|---|
| Customer value | Would a customer agree a higher number means they're better off | Metric rewards frequency of visits, not depth of use |
| Leading indicator | Moves before revenue/retention, not after | Metric only confirms outcomes that already happened |
| Gaming resistance | Can't rise without real value also rising | A single lever (notifications, defaults, friction removal) moves it in isolation |
The DAU Trap: When a "Good-Looking" Metric Wrecks the Product
Daily Active Users is the textbook cautionary tale because it looks unambiguously good and is unambiguously gameable — a team can inflate it through notification volume, autoplay, or dark-pattern re-engagement loops without a single customer getting more value. The lesson isn't "never use DAU" — it's that any single-number North Star needs a paired guardrail metric, or the incentive system optimizes for the easiest lever, not the intended one.
This isn't a hypothetical. In January 2018, Facebook publicly announced it would deprioritize pure time-spent as a goal for its News Feed ranking, after its own research reportedly found that passive content consumption correlated with worse self-reported well-being than active social interaction — while active interaction correlated with better outcomes. The company had spent years optimizing engagement metrics that looked healthy in aggregate while masking a real difference in the kind of engagement being produced. That gap between "the number is up" and "customers are better off" is exactly what a North Star test is supposed to catch before it ships, not after a public course correction.
This is precisely the failure mode economist Charles Goodhart described decades before growth teams existed: "When a measure becomes a target, it ceases to be a good measure."* Goodhart's Law isn't a growth-specific idea — it's a general property of any system where people know they're being measured and have discretion over how to respond. A North Star metric is a target by definition, which means it's structurally exposed to Goodhart's Law from the day it's chosen, not just after a team starts "gaming" it on purpose.
| Signal | Looks like growth | Actually measures |
|---|---|---|
DAU | More daily visits | Notification effectiveness, not value delivered |
| Signups | Funnel growth | Ease of the form, not fit with the product |
| Session length | Deeper engagement | Confusing navigation as easily as genuine interest |
| Feature adoption clicks | Feature-market fit | UI discoverability, not whether the feature solved anything |
The pattern across all four rows: each metric can be moved by a single, cheap lever that has nothing to do with customer value. That's the tell a stress-test checklist is built to catch early.
A Stress-Test Checklist: Gaming-Proofing Your Candidate Metric
Before anchoring a roadmap to a candidate metric, run it through a structured stress test rather than a gut check — gaming vectors are rarely obvious until you deliberately hunt for them. The checklist below is designed to be run in a single working session with your growth and analytics leads in the room.
- List every cheap lever. For each way the number could move without value moving (more notifications, lower friction, gamed defaults), write it down explicitly — don't stop at the first one you find.
- Pair it with a guardrail metric. Choose a second number that would visibly suffer if the North Star were gamed — for
DAU, a guardrail might be session quality or unsubscribe rate. - Ask what a bad actor on the team would do. If a PM were rewarded purely on this number with no other oversight, what's the laziest way to move it? If the answer is uncomfortable, the metric needs a constraint.
- Check for a single point of control. If one team, one feature, or one lever can move the metric in isolation, it's too narrow — a durable North Star usually requires cross-functional effort to move.
- Trace it to a
Jobs to Be Doneoutcome. Confirm the metric only rises when a customer's job gets done better, not merely attempted more often. - Run a small experiment first. Before committing org-wide, test whether the metric can be moved artificially in a contained experiment — experiment velocity practices matter here because a fast test-and-learn cadence is what lets you catch a gameable metric in weeks instead of discovering it a year into a bad roadmap.
- Re-test after six months. A metric that was hard to game at launch can become gameable once the team learns its shortcuts — schedule a re-audit, don't assume day-one integrity holds forever.
Alistair Croll and Benjamin Yoskovitz, authors of Lean Analytics, popularized the idea of the One Metric That Matters (OMTM) — a single number that changes based on what stage a business is actually in. Their framework's implicit warning is as important as its structure: the right metric for seed-stage product-market fit is often actively wrong once you're scaling, because the gaming vectors and value signals shift underneath it.
Trace the Ripple Effects Before You Commit
A North Star metric doesn't operate in isolation — it sits inside a causal system where pushing it changes incentives for adjacent teams, and some of those downstream effects are perverse ones nobody intended. The fastest way to catch a perverse incentive is to trace the metric's ripple effects on paper — or in a model — before a single sprint gets planned around it, not after a quarter of results comes back strange.
Systems thinker Donella Meadows, in Thinking in Systems: A Primer, described goals and metrics as some of the most powerful — and most dangerous — leverage points in any system, precisely because everything downstream reorganizes around them. A North Star metric is that kind of leverage point: change it and you've changed what dozens of individual decisions will optimize for, whether or not that's what you modeled.
This is also where North Star selection connects to funnel ownership. As the growth PM funnel ownership and leverage guide covers, a metric chosen without mapping its position in the funnel tends to reward whichever stage is cheapest to move, not the stage with the most real leverage — usually the earliest, easiest-to-juice step rather than the one that actually predicts retention.
Amplitude's widely-referenced North Star Playbook — associated with product coach John Cutler — frames this same idea as an "input metrics" tree feeding the North Star from below. The practical takeaway: don't just pick the top-line number, map the two or three inputs that feed it, because that's where a gaming vector or a perverse incentive is actually visible.
Choosing Between Candidate Metrics: A Side-by-Side Framework
When you have two or three plausible North Star candidates, score them side by side against the same three criteria rather than debating in the abstract — a structured comparison surfaces trade-offs that pure discussion tends to paper over. Below is a worked example for a subscription marketplace product choosing between three real candidates.
| Candidate metric | Reflects customer value | Leading indicator | Gaming resistance | Verdict |
|---|---|---|---|---|
| Weekly active buyers | Medium — activity, not outcome | Strong — precedes revenue | Medium — gameable via reminder emails | Guardrail metric, not North Star |
| Completed transactions per active buyer | Strong — ties directly to job completion | Strong — precedes repeat revenue | Strong — hard to fake a completed transaction | Best North Star candidate |
| Total signups | Weak — measures acquisition, not use | Weak — lags true adoption | Weak — cheap to inflate via paid spend | Reject as North Star |
The table makes the decision almost mechanical once the scoring is honest: completed transactions per active buyer wins because it can't rise without a real job getting done, it moves ahead of revenue, and there's no single cheap lever that fakes a completed transaction the way there is for a signup or a login. Weekly active buyers survives as a useful guardrail metric precisely because it's a decent leading signal but too gameable to anchor the whole team to.
Run this same table with your own candidates before you commit publicly to one number — a scoring exercise that takes an afternoon is considerably cheaper than a year spent optimizing the wrong incentive.
Key Takeaways
- A North Star metric is an incentive system, not a dashboard number — every team downstream will find the cheapest way to move it, whether or not that path creates real value.
- Three tests decide fitness: the metric reflects real customer value, moves earlier than lagging outcomes like revenue, and resists being inflated without value also rising.
DAUis the classic cautionary tale, not because it's a bad number in isolation, but because it's easy to move through notifications and re-engagement tricks with zero value delivered.- Goodhart's Law is not optional — any metric turned into a target is structurally exposed to being gamed, which is why a stress-test checklist has to be run before rollout, not after.
- Pair every North Star with a guardrail metric that would visibly degrade if the North Star were gamed, so a cheap win on paper shows up as a real cost somewhere else.
- Trace ripple effects before committing — a candidate metric's downstream incentives are often easier to see in a causal model than in a spreadsheet debate.
- Re-audit every six months — a metric that was hard to game at launch can develop gaming vectors once the team learns the system's shortcuts.
Frequently Asked Questions
What is a north star metric?
A North Star metric is the single number a product team organizes roadmap and prioritization decisions around, chosen because it's believed to causally predict long-term business success through genuine customer value — not just a KPI on a dashboard. It differs from a lagging business metric like revenue because it's meant to move first.
How is a north star metric different from OKRs?
A North Star metric is one durable number that rarely changes; OKRs are typically quarterly goals and initiatives that should ladder up to moving that North Star. Treat the North Star as the constant and OKRs as the rotating set of bets a team makes each quarter to move it.
Can a company have more than one north star metric?
Most frameworks recommend exactly one, paired with two or three guardrail metrics, because splitting focus across multiple "North Stars" recreates the coordination problem a single metric was meant to solve. Different teams or business units can have their own North Stars, but each team should have only one.
How often should you revisit your north star metric?
Re-audit a North Star roughly every six to twelve months, or immediately after a major business model shift, because gaming vectors that didn't exist at launch tend to emerge once a team has learned the system's shortcuts. A metric review shouldn't be a one-time exercise at launch.
What's an example of a good north star metric?
A commonly cited example is Airbnb's historical focus on nights booked rather than signups or listing counts, because it only rises when a guest and host both complete a real transaction — a structure similar to the completed transactions per active buyer example above. The specific number matters less than whether it passes the value, leading-indicator, and gaming-resistance tests for your own product.