Success in a revenue-driven product is a curve that goes up and to the right. Success in government is a specific person receiving the benefit, protection, or service they were legally owed — on time, without having to fight for it. Public-value metrics measure that outcome directly, not the paperwork volume surrounding it.
Public-value metrics track whether people who were entitled to a benefit or service actually received it — not how many applications your team processed. Build a logic model that traces outputs to outcomes, disaggregate results by group so averages can't hide failure, and price
cost-per-successful-outcomeinstead of cost-per-transaction.
For a PM used to treating non-revenue product metrics as a rare edge case — the free tier, the internal tool, the nonprofit side project — government is where they're the only kind that exists. There is no MRR to eventually punish a bad metric. That means the discipline of measuring government outcomes has to be built in deliberately, not discovered later when a dashboard number stops matching reality on the ground.
Why Application Counts Aren't Public Value
Applications submitted, calls answered, and tickets closed are outputs — proof your team did something, not proof anyone's situation improved. They're popular because they're cheap to count and hard to argue with in a budget hearing, but an agency can hit every output target and still fail the people the program exists for.
This is the same trap product teams eventually learn to avoid with MAU or page views — a number that moves without telling you whether it moved for a good reason. Government just has no revenue signal to eventually force the correction. A private company that optimizes the wrong metric loses customers and, eventually, revenue. A public program that optimizes the wrong metric can run for years with a clean quarterly report and a growing population it's quietly failing.
Mark Friedman's Results-Based Accountability (RBA) framework, widely used across US state and local government, draws this line explicitly: performance measures ("how much did we do," "how well did we do it") are not the same as population or outcome measures ("is anyone better off"). Most agency dashboards are stacked with the first kind. Almost none default to the second, and almost nobody asks why.
The mismatch shows up the same way across very different services:
- Benefits processing — applications submitted or processed, versus eligible people who actually receive their correct benefit payment.
- Permitting — permits issued, versus businesses that opened and stayed open twelve months later.
- Case management — cases opened or closed, versus families who exited the system stable rather than returning within a year.
- Public health outreach — flyers distributed or clinics held, versus vaccination or screening rates in the target population.
- Workforce and education programs — workshops delivered, versus participants who gained the skill and used it to get or keep a job.
None of this means outputs are worthless — a program with zero output can't produce any outcome, so outputs remain a legitimate leading indicator. The distinction matters because outputs are where public-sector teams default to reporting, especially inside the legacy procurement and reporting cycles covered in the broader govtech product management playbook. Outcomes are where accountability actually lives, and they're the harder, less comfortable number to publish.
From Outputs to Outcomes: Building a Logic Model You Can Defend
A logic model turns a vague claim like "this program helps people" into a traceable chain: inputs fund activities, activities produce outputs, outputs should drive outcomes, and outcomes accumulate into impact. Each arrow in that chain is a claim you can test — and most programs have never actually tested theirs.
The W.K. Kellogg Foundation Logic Model Development Guide is the reference most public-sector evaluators reach for, and it's useful precisely because it forces the five links to be written down separately instead of blurred into one paragraph of program description. Skipping a link is how most metrics quietly become vanity metrics — the output gets measured because it's the easiest link to see, and the outcome gets assumed rather than tested.
The table below walks one common service — unemployment insurance claims — through all five stages, since it's a program almost every US state runs and almost every state reports on by output alone.
| Logic model stage | Definition | Unemployment insurance example |
|---|---|---|
| Input | Resources committed | Staff, claims software, benefit trust fund |
| Activity | What the program does | Process claims, verify eligibility, issue payments |
| Output | Volume of activity produced | Claims processed within a target number of days |
| Outcome | Change for the person | Eligible claimant receives the correct payment before financial hardship hits |
| Impact | Population-level change | Reduced local poverty and eviction rates during unemployment spells |
Notice that "claims processed within a target number of days" sits at the output stage, not the outcome stage. A state can hit that target while still paying the wrong amount, denying eligible claimants, or approving people who no longer qualify. This is the mechanical core of measuring government outcomes instead of government activity: asking whether the right payment reached the right person in time to matter, not just whether the queue moved.
A logic model shows the intended chain. A causal loop shows what actually happens once people start optimizing for the metric you chose, including the loops nobody designed on purpose. Take the same claims example: a state sets a fixed-day processing target as its output goal.
Caseworkers under pressure to hit that target start approving borderline claims faster and scrutinizing edge cases less — a reinforcing loop where faster processing looks good on the dashboard. Denial-in-error and approval-in-error rates both climb quietly underneath it. Appeals volume rises months later, consuming the same caseworker hours the speed target was supposed to free up — a balancing loop where the system fights itself. The output metric can move in the right direction for a full year while the outcome gets worse underneath it.
The Perverse-Incentive Trap: When a Proxy Metric Drives Bad Behavior
Any metric that matters enough for careers, budgets, or headlines eventually gets optimized directly — and the moment that happens, it stops measuring what it was built to measure. This isn't hypothetical. It has a name, decades of public-administration documentation behind it, and a track record of happening to metrics that looked exactly as reasonable as the one you're about to adopt.
Economist Charles Goodhart observed that once a measure becomes a target, it ceases to be a good measure — widely shortened to Goodhart's Law. Social scientist Donald T. Campbell reached the same conclusion independently while studying US social-program evaluation in the 1970s, and stated it more bluntly.
"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." — Donald T. Campbell, 1976
The clearest modern illustration is the US Department of Veterans Affairs wait-time scandal, made public in 2014. VA leadership had set a 14-day scheduling target, and that target fed directly into facility bonuses and executive performance reviews. The metric was simple, trackable, and politically legible — everything a good proxy metric is supposed to be.
A subsequent VA Office of Inspector General investigation found that schedulers at multiple facilities, including the Phoenix VA, had kept secret, unofficial waiting lists. The official wait-time numbers stayed compliant while real veterans, some seriously ill, waited months longer than the reported figures suggested. The scandal cost the VA Secretary his job and triggered a nationwide audit, but the underlying lesson generalizes to any program that lets frontline staff influence the numerator and denominator of the metric they're judged on.
The same pattern surfaced in US education a few years earlier. No Child Left Behind tied school funding and staff evaluations to standardized test pass rates — a reasonable-sounding proxy for "students are learning." The 2011 Atlanta Public Schools cheating investigation found systematic, coordinated correcting of student answer sheets across dozens of schools, driven by the same pressure: the test score became the thing being managed, not the learning it was meant to represent.
Before adopting any proxy metric for a public program, run it through the same four questions:
- Who controls the inputs to this number, and could they benefit from moving it without moving the outcome?
- What's the easiest way to game this metric without breaking any explicit rule?
- Does the metric reward speed or volume over correctness — and can those trade off against each other?
- Would this number look the same if the underlying outcome got worse? If yes, it's a vanity metric wearing an outcome costume.
Measuring What Matters: Equity-Disaggregated Rates and Cost-Per-Successful-Outcome
A single blended average can hide total failure for a subgroup behind overall success — the fix is disaggregating every outcome by the groups most likely to be underserved. The second fix is pricing programs by cost-per-successful-outcome, not cost-per-application, so cheap-but-ineffective delivery stops looking efficient on a budget slide.
Equity-disaggregated metrics
A benefits program that serves 92% of eligible applicants overall can still be failing non-English speakers, rural residents, or applicants with disabilities almost entirely — a single blended number will never show it. Disaggregation means reporting the same outcome rate separately by race, income band, geography, language, disability status, and age, then watching the gaps between groups, not just the overall trend.
Initiatives like What Works Cities (Bloomberg Philanthropies' program working with hundreds of US municipalities since 2015) and Results for America have spent a decade pushing local governments toward exactly this practice — outcome measurement that's disaggregated by default, not added later as an afterthought once someone complains.
Disability status deserves its own line in that breakdown, not a footnote. A digital service can clear the baseline legal bar covered in the accessibility legal baseline under Section 508 and still produce a materially worse outcome rate for people using assistive technology. Compliance and outcome equity are related but not the same test, and only disaggregated outcome data catches the gap between them.
Cost-per-successful-outcome
Most public budgets get justified on cost-per-output — cost per application processed, cost per call handled, cost per case opened. That number rewards throughput regardless of whether the throughput helped anyone, and it can make a program that quickly rejects or churns people look more "efficient" than one that works harder to get them approved.
cost-per-successful-outcome divides total program cost by the number of people who received the outcome they were entitled to — not the number processed, approved on paper, or merely contacted. It's a harder number to produce, since it requires knowing not just what you output but what stuck, and that's exactly why it's more honest.
The UK's HM Treasury Magenta Book and Green Book — the government's own guidance for evaluating and appraising public spending — both push agencies toward outcome-based cost-effectiveness analysis over simple unit-cost-per-transaction, for the same reason: a program can look cheap per transaction while being expensive per person actually helped, once you count the rework, appeals, and re-applications a bad process generates.
The gap between these two numbers is usually where a program's real performance — good or bad — is hiding, as the comparison below shows across three common service types.
| Program | Cost-per-output (common) | Cost-per-successful-outcome (harder, more honest) | What the gap usually hides |
|---|---|---|---|
| Benefits eligibility screening | Cost per application processed | Cost per eligible person who receives correct, timely payment | Wrongful denials, appeal rework, re-applications |
| Small-business permitting | Cost per permit issued | Cost per business that opens and survives 12 months | Permits issued to businesses that never open |
| Workforce training program | Cost per participant enrolled | Cost per participant who gains and keeps a job | High enrollment, low completion, no placement follow-up |
In every row, the honest number requires tracking something past the point where your team's direct responsibility ends — whether a business survived, whether a payment was correct, whether a job stuck. That's exactly the data most legacy case-management systems were never built to capture, which is where instrumentation becomes the real bottleneck.
Instrumenting the Chain: Tracing Outputs to Outcomes in Practice
Measuring public value requires data your system probably doesn't collect today — outcome data almost always lives downstream of the transaction your platform actually owns, often in a different agency's system entirely. Closing that gap starts with naming the outcome before you build anything, not after.
Most legacy government systems were built to record transactions — an application submitted, a payment issued, a case closed — because that's what the original program needed to function. They were never instrumented to answer "and then what happened to this person six months later," which is precisely the outcome question. Closing that gap is one of the harder, less glamorous parts of the legacy modernization strategy public-sector teams have to plan for, since it often means integrating with a labor department, a health system, or a court's data, not just improving your own frontend.
Before you can pick the right outcome metric, you have to be precise about the job the person is actually hiring the service to do, not the job your agency assumes it's providing. A framework like the one covered in the complete guide to jobs-to-be-done is as useful in govtech as in any commercial product: a citizen doesn't want a permit, they want a legally open business; they don't want a processed claim, they want their rent paid on time. Define the job first, and the outcome metric follows almost automatically.
The other instrumentation gap is journey-shaped, not data-shaped: most agencies can see the touchpoint they own and almost nothing before or after it. Mapping the full applicant journey, including the emotional and practical friction on either side of your service's part in it, is what surfaces the silent drop-off between "submitted an application" and "received the benefit" — the exact gap a pure output metric can never see.
None of this instrumentation happens by accident inside a government contract — it has to be specified. If a vendor's statement of work only pays for output SLAs ("process X applications per month"), that's the only thing the vendor has any incentive to optimize. Outcome-based contracting has to be written into requirements from the start, which is exactly the kind of tradeoff covered in designing a product around government procurement constraints — the metric you procure for is usually the metric you get.
Where a causal-loop view helps you interrogate the chain
This is the specific gap Prodinja's Systems Engineering studio is built around: causal-loop diagrams that let you map how an output — claims processed within a target window, permits issued, workshops delivered — actually connects, or doesn't, to the outcome you're claiming credit for, including reinforcing and balancing loops like the caseworker example above. The prototype's Ask your data view is designed to let you interrogate that chain directly, asking whether the output you track is genuinely wired to the outcome you report rather than assuming the arrow between them holds. It's built as a way to walk through your own causal assumptions, not as an oracle that hands you a verified answer.
Key Takeaways
- Outputs measure activity; outcomes measure whether a person's situation actually changed — applications processed, calls answered, and permits issued are leading indicators at best, not proof of public value.
- Build a logic model (input → activity → output → outcome → impact) before you pick a metric, using a structure like the
W.K. Kellogg Foundation's logic model guide, so every arrow in the chain is a claim you can test rather than assume. - Any metric intense enough to affect careers or budgets will eventually be gamed — Goodhart's Law and Campbell's Law aren't theoretical, and the 2014 VA wait-time scandal shows exactly how a reasonable-looking proxy metric produced falsified data and real harm.
- Disaggregate every outcome by race, income, geography, disability, and language before trusting a blended average — a program can look successful overall while failing a specific group almost entirely.
- Price programs by
cost-per-successful-outcome, not cost-per-transaction — a cheap-per-application program that generates appeals, rework, and re-applications is usually more expensive per person actually helped than it looks. - Instrumentation is usually the real blocker, not analysis — outcome data typically lives downstream of your system, in another agency's records, which makes data-sharing and journey-mapping as important as the metric definition itself.
- Write the outcome metric into procurement and contracts, not just the dashboard — a vendor paid against output SLAs will optimize outputs, no matter what your team privately cares about.
Frequently Asked Questions
What's the difference between an output metric and an outcome metric in government programs?
An output metric counts what your team produced — applications processed, permits issued, calls answered. An outcome metric measures what changed for the person the program serves — whether they received the benefit, kept their job, or stayed housed. A program can hit every output target while its outcome quietly gets worse, which is why the two need to be tracked and reported separately.
How do you measure success in public-sector product management without a revenue metric?
Replace revenue with a named public-value outcome and a logic model connecting your output to it, defined before launch rather than reverse-engineered from whatever data happens to exist. Borrow Results-Based Accountability's discipline of separating "how much did we do" from "is anyone better off," and report both — eliminating the first isn't the goal, conflating it with the second is the mistake.
What is cost-per-successful-outcome and how is it calculated?
cost-per-successful-outcome divides total program cost by the number of people who actually received the outcome they were entitled to, not the number processed or contacted. It's calculated the same way a unit-economics metric is in commercial product work, just with "successful outcome" substituted for "paying customer" as the denominator's qualifying event.
Why do government performance metrics get gamed so often?
Because any metric tied to funding, bonuses, or headlines creates pressure to move the number directly, and frontline staff often have some control over how it's counted — this is Goodhart's Law and Campbell's Law in practice, not a failure specific to any one agency. The 2014 VA wait-time scandal and the 2011 Atlanta Public Schools testing scandal are two well-documented, independent examples of the same underlying dynamic.
How do you disaggregate outcome metrics without creating an unmanageable number of dashboards?
Pick the two or three disaggregation dimensions with the most documented equity risk for your specific program, commonly race, income, disability status, and geography, rather than slicing every possible variable. Initiatives like What Works Cities and Results for America have published disaggregation playbooks that most local governments can adapt instead of building the taxonomy from scratch.