Citizens can't opt out of government the way a customer opts out of a loyalty program, so filing taxes or applying for benefits is mandatory, not chosen. That asymmetry redefines the PM's job: you're not building a data asset, you're running a custody chain — collect only what law requires, use it only as authorized, delete it on schedule.
Quick answer: Government PMs are custodians of citizen data under legal mandate, not owners building a growth asset. Practice purpose limitation, enforce retention schedules, and separate "can access" from "authorized to use" before every data decision.
Custodian, Not Owner: Why the Mental Model Matters
A custodian holds something on behalf of someone else, bound by rules they didn't write and can't unilaterally change. That's a public-sector PM's actual relationship to citizen data — closer to a trustee than a growth operator with a data warehouse to fill.
Commercial product management treats data as an asset to accumulate: more signal, better personalization, richer models. Public-sector data work runs on the opposite instinct. Every field you collect is a liability until proven necessary, because the citizen in front of you had no meaningful choice about interacting with your agency in the first place.
Two academic frameworks make this concrete instead of just aspirational:
- Daniel Solove's taxonomy of privacy (2006) breaks privacy harm into concrete categories — surveillance, aggregation, secondary use, exclusion — rather than one vague "privacy" bucket. It's a useful checklist when a feature review gets stuck on "is this a privacy issue?"
- Helen Nissenbaum's contextual integrity framework argues data flows are only appropriate within the norms of the context they were collected in. A citizen's disability status shared with a benefits office is normal; the same field flowing to a licensing board is a context violation, even with the same legal owner of the data.
Both frameworks point at the same discipline: the question isn't "do we have this data," it's "is this specific use consistent with why the citizen gave it to us." If you're new to the broader discipline of building government software, the complete guide to govtech product management is a good starting point — this piece goes deep on the one slice of it that carries the most legal exposure.
The Legal Mandates That Define Your Boundaries
Public-sector data governance isn't a values statement your team writes in a workshop — it's a set of external, binding obligations that predate your product and will outlast your tenure. Four anchor frameworks show up in some form in nearly every jurisdiction.
| Framework | Origin | Core Principle | What It Means for Your Backlog |
|---|---|---|---|
Privacy Act of 1974 | US federal law | Systems of Records must be published and use limited to the stated purpose | A new data store needs a purpose statement before code is written, not after |
GDPR Article 5 | EU law, widely copied by state and sector-specific US rules | Purpose limitation, data minimization, storage limitation | Fields default to "not collected" unless mapped to a lawful basis |
OECD Privacy Guidelines (1980, revised 2013) | International, adopted in some form by dozens of countries | Collection limitation, use limitation, accountability | A baseline checklist for any new intake form or API |
NIST Privacy Framework (2020) | US, voluntary but widely adopted by federal and state agencies | Risk-based privacy controls tied to program outcomes | Shared vocabulary with your security and compliance partners |
None of these frameworks were written with product velocity in mind, and that's the point. A Privacy Act System of Records Notice, once filed, effectively becomes a spec — it tells you what data you're allowed to hold and why, and amending it is slower than shipping a feature flag.
Much of this exposure gets locked in before a single screen gets designed, often inside the same conversations that shape how government procurement rules shape what you can build. A data-sharing clause buried in an RFP or an interagency agreement can commit your product to a retention policy months before product management is even in the room.
Privacy sits alongside a small set of other legal non-negotiables in this work, the same way Section 508 accessibility compliance functions as a legal floor, not a nice-to-have. Neither is a backlog item you deprioritize under deadline pressure — both are conditions of the product existing at all.
Purpose Limitation: The Question Every Feature Must Answer
Purpose limitation means data collected for one stated reason cannot silently migrate to serve a different one, even inside the same agency. It's the single principle most responsible for the "wait, we can't do that with this data" conversations that stall roadmaps late in a sprint.
The failure mode is rarely malicious. It's a fraud-detection field that eligibility staff start using for general case triage. It's a phone number collected for appointment reminders that ends up feeding an outbound enforcement campaign. Each reuse feels incremental; none of them were authorized by the original collection notice.
A borrowed discipline helps here more than a legal checklist does. Treat every data field like a hired employee and ask what job it's actually hired to do — the same interrogation taught in the complete guide to Jobs to Be Done. If you can't name the specific job a field performs in the specific feature you're building, you can't justify collecting or reusing it.
Run this test before any feature that touches an existing field for a new purpose:
- What was the original stated purpose at collection, per the notice or consent language?
- Is the new use a natural extension of that purpose, or a different program entirely?
- Does a statute, regulation, or interagency agreement explicitly authorize the new use — not "nothing prohibits it," but an affirmative basis?
- Would the citizen who provided this data reasonably expect this new use, given the context they gave it in (Nissenbaum's test again)?
- If the answer to #3 is unclear, who owns the decision — legal, privacy office, program leadership — and have you actually routed it there?
If any answer is "no" or "not sure," the feature needs a new legal basis before it needs a design review.
Retention Schedules: Data Has an Expiration Date
Retention is the most neglected half of data governance because deletion has no visible feature and no demo. Nobody celebrates a records-disposal job running successfully; the failure mode — a decade-old dataset nobody remembers the justification for — is invisible until a breach, audit, or FOIA request finds it.
Records-management authorities like the National Archives and Records Administration (NARA) publish General Records Schedules precisely because agencies otherwise default to keeping everything indefinitely. The pattern below is illustrative, not a universal rule — actual windows vary by program and jurisdiction, and your agency's records officer has the authoritative schedule.
| Data Category | Typical Legal Basis | Illustrative Retention Pattern | Disposal Trigger |
|---|---|---|---|
| Identity verification documents | Program eligibility statute | Life of case plus a fixed post-closure window | Case closure and the statutory window elapses |
| Financial or income records | Benefits or tax administration statute | Multi-year, tied to the agency's records schedule | Determination finalized and the appeal window closes |
| Communication logs (calls, chat, email) | Service-delivery policy | Commonly shorter than the underlying case file | Case closure or a defined service-log retention policy |
| Fraud investigation records | Inspector-general or law-enforcement authority | A separate, often longer schedule than the case itself | Investigation closed and any statute of limitations passes |
| Aggregate or de-identified analytics data | Program-evaluation authority | Can extend past individual case retention once properly de-identified | Not applicable once de-identification is verified and re-identification risk is reviewed |
Read across the table and one thing jumps out: retention isn't one number, it's a portfolio of different clocks running against different legal bases inside the same case. A product that treats "the case record" as a single retention unit will either delete evidence it's still legally required to hold, or keep communication logs years past their own policy window.
Legacy systems are where these mismatched clocks accumulate for decades — fields nobody can trace back to a statute, retention jobs that were never built, schedules that predate the system itself. That's exactly the terrain a public-sector legacy modernization strategy has to reckon with before any migration, because you cannot migrate a retention obligation you can't first identify.
Treat a missing retention trigger the same way you'd treat a missing acceptance criterion: the feature isn't done until deletion has an owner and a date, not just collection.
Access vs. Authorization: The Line Most PMs Miss
Technical access and legal authorization are two different questions, and conflating them is where most well-intentioned governance programs break down. A caseworker's role permissions might let them query every program a citizen has ever touched; the statute governing their specific program almost never authorizes using that full view.
This distinction rarely shows up in a data model, because access control and legal authorization live in different systems entirely — one in your identity and access management (IAM) layer, one in a policy document your engineering team has probably never read.
| Scenario | Can They Access It? | Are They Authorized to Use It? |
|---|---|---|
| Caseworker views cross-program case history via an admin query tool | Yes — role permits broad read access | Usually no — statute scopes use to the specific program they administer |
| Fraud analyst reuses a benefits application field for an unrelated enforcement lead | Yes — same database, same schema | No, absent an explicit interagency data-sharing agreement |
| Support agent reads a citizen's full communication log to resolve one ticket | Yes — support tooling grants full history | Partially — only the portion relevant to the current service request |
| Analytics team joins two programs' datasets on a shared citizen identifier | Yes — identifiers match, join is technically trivial | No, unless both programs' authorizing statutes explicitly permit the linkage |
Notice the pattern: the technical answer is almost always "yes," and the legal answer is almost always "it depends on a document your product team hasn't read." That gap is exactly where PMs need to build friction into the product deliberately — narrower default views, purpose-tagged queries, audit logs that capture why a record was accessed, not just that it was.
- Design defaults for the narrowest authorized view, and require an explicit, logged reason to expand it — never the reverse.
- Tag data access requests with a stated purpose at the point of query, so an audit can reconstruct intent later, not just activity.
- Treat "the admin tool lets you see everything" as a bug report, not a feature — broad internal visibility is usually an accident of shared infrastructure, not a deliberate authorization decision.
The Entity-Modeling Exercise That Exposes Over-Collection
Naming every entity and relationship in a citizen dataset — out loud, in a diagram, before writing a single query — reliably surfaces over-collection that a requirements document never would. Fields hide inside vague entity names; once you have to label the entity precisely, the unjustified fields have nowhere left to hide.
Take a typical benefits-eligibility case and sketch its entities:
- Citizen (the applicant)
- Household (linked citizens sharing an eligibility determination)
- Application (a single submission, tied to a household)
- Eligibility Determination (a decision, referencing verified evidence)
- Benefit Award (tied to a determination)
- Payment (tied to an award)
- Case Note (attached to an application)
- Communication Record (tied to the citizen, across channels)
- Fraud Flag (often, tellingly, tied to the citizen — not the case)
Two entities in that list are where over-collection hides in plain sight. The Citizen entity is where every field ever requested across every program tends to accumulate — a biometric photo captured for one pilot, an immigration-status field needed by exactly one benefit type, a veteran-status flag nobody remembers assigning ownership of. The Fraud Flag entity is worse: it's frequently modeled as citizen-scoped rather than case-scoped, which quietly turns a single program's suspicion into a cross-program blacklist nobody authorized.
The exercise works because modeling forces specificity that a feature request never does. "We need identity info" survives a planning meeting. "We need a national_id_number column on the shared Citizen table, readable by every program's caseworkers" does not — once it's written down that plainly, someone in the room asks why.
Mapping where each data point actually enters the system across a citizen's end-to-end journey usually reinforces the same finding: the same income document gets requested at three separate touchpoints because no single entity owns the authoritative copy, and each intake team quietly re-collects rather than re-using what already exists.
A Data-Minimization Decision Framework
Once the entity model is on the table, run every field on the Citizen and Household entities through this six-question test before it ships:
- Necessity: Is this field required to make the specific determination this feature performs — not "might be useful someday"?
- Least-identifying proxy: Could a category, a range, or a token substitute for the raw identifier and still satisfy #1?
- Purpose lineage: Which statute or program explicitly authorizes collecting this field, and does that authorization survive if the record is later joined with another program's data?
- Downstream exposure: Which other systems, reports, or integrations will touch this field — and does each one have its own lawful basis, rather than inheriting one from the original collection?
- Retention trigger: What event ends the need for this field, and is a disposal job actually scheduled against that trigger, or just assumed?
- Re-identification risk: Combined with the rest of the record, does this field single out an individual in a way disproportionate to the program's actual purpose?
A field that fails #1 gets cut. A field that fails #3 or #5 gets escalated to legal or the privacy office before it ships — not after an audit finds it.
Where a Data Model Makes This Governance Concrete
Everything above is easier to say than to do inside a real product, because most teams never actually draw the entity diagram — they infer it from whatever the database already looks like. Prodinja's Data Modelling studio is built around forcing that step explicitly, before a single table exists.
You name every entity, every relationship, and every field in a citizen dataset before generating schema. That's designed to surface exactly the kind of citizen-scoped fraud flag or all-programs identifier described above while it's still a diagram, not a migration.
It's a modeling discipline, not a compliance verdict — the legal judgment calls in the framework above still belong to your privacy office. But naming what you're collecting is the prerequisite for deciding what you're actually allowed to keep.
Key Takeaways
- You're a custodian, not an owner — citizen data is held under legal mandate, not accumulated as a growth asset, because citizens can't opt out of the interaction that generated it.
- Anchor every data decision to a named legal basis —
Privacy Act of 1974-style systems-of-records notices,GDPR Article 5,OECD Privacy Guidelines, or your jurisdiction's equivalent — not to "nothing prohibits it." - Purpose limitation means one field, one job — if you can't name the specific authorized purpose a field serves in this feature, it isn't ready to ship.
- Retention is a portfolio of clocks, not one date — identity documents, financial records, communication logs, and fraud files each run on different legal timelines inside the same case.
- Access and authorization are different questions — broad technical visibility is usually an infrastructure accident, not a deliberate, legally reviewed grant of authority.
- Entity modeling is a governance tool, not just a technical one — naming a field precisely on a diagram exposes over-collection that a requirements doc lets slide.
Frequently Asked Questions
What is the difference between data privacy and data governance in government products?
Data privacy sets citizen-facing limits on collection and use; data governance is the internal system of ownership, retention rules, and access controls that enforces those limits day to day. Privacy answers "should we collect this"; governance answers "who owns it, how long do we keep it, and who's authorized to touch it."
How does GDPR apply to US public-sector agencies?
Most US public-sector agencies aren't directly bound by GDPR unless they process EU residents' data, but its principles — purpose limitation, data minimization, storage limitation — are widely mirrored in US state privacy laws and federal guidance like the NIST Privacy Framework. Treat GDPR's structure as a well-tested reference model even where it isn't the controlling law.
What is purpose limitation in plain terms?
Purpose limitation means data collected for one stated reason cannot be reused for a materially different reason without new, explicit legal authorization. A phone number collected for appointment reminders, for example, isn't automatically fair game for an unrelated enforcement outreach campaign.
How long should a government agency keep citizen data?
There's no single answer — retention windows are set by program-specific statutes and your agency's records schedule (often issued or modeled on NARA guidance), and different data types inside the same case can carry different retention periods. Identity documents, financial records, and communication logs commonly follow different clocks even within one case file.
Who is accountable when citizen data is over-collected or misused?
Accountability typically sits with the agency as data controller under its governing statute, but product management carries real responsibility for the fields and access patterns actually built into the system. Government Accountability Office (GAO) reviews have repeatedly found agencies with gaps in privacy or records-retention controls, which is exactly the exposure a disciplined entity-modeling and purpose-limitation practice is designed to close.