Citizens can't opt out of government the way a customer opts out of a loyalty program, so filing taxes or applying for benefits is mandatory, not chosen. That asymmetry redefines the PM's job: you're not building a data asset, you're running a custody chain — collect only what law requires, use it only as authorized, delete it on schedule.

Quick answer: Government PMs are custodians of citizen data under legal mandate, not owners building a growth asset. Practice purpose limitation, enforce retention schedules, and separate "can access" from "authorized to use" before every data decision.

Custodian, Not Owner: Why the Mental Model Matters

A custodian holds something on behalf of someone else, bound by rules they didn't write and can't unilaterally change. That's a public-sector PM's actual relationship to citizen data — closer to a trustee than a growth operator with a data warehouse to fill.

Commercial product management treats data as an asset to accumulate: more signal, better personalization, richer models. Public-sector data work runs on the opposite instinct. Every field you collect is a liability until proven necessary, because the citizen in front of you had no meaningful choice about interacting with your agency in the first place.

Two academic frameworks make this concrete instead of just aspirational:

  • Daniel Solove's taxonomy of privacy (2006) breaks privacy harm into concrete categories — surveillance, aggregation, secondary use, exclusion — rather than one vague "privacy" bucket. It's a useful checklist when a feature review gets stuck on "is this a privacy issue?"
  • Helen Nissenbaum's contextual integrity framework argues data flows are only appropriate within the norms of the context they were collected in. A citizen's disability status shared with a benefits office is normal; the same field flowing to a licensing board is a context violation, even with the same legal owner of the data.

Both frameworks point at the same discipline: the question isn't "do we have this data," it's "is this specific use consistent with why the citizen gave it to us." If you're new to the broader discipline of building government software, the complete guide to govtech product management is a good starting point — this piece goes deep on the one slice of it that carries the most legal exposure.

Public-sector data governance isn't a values statement your team writes in a workshop — it's a set of external, binding obligations that predate your product and will outlast your tenure. Four anchor frameworks show up in some form in nearly every jurisdiction.

FrameworkOriginCore PrincipleWhat It Means for Your Backlog
Privacy Act of 1974US federal lawSystems of Records must be published and use limited to the stated purposeA new data store needs a purpose statement before code is written, not after
GDPR Article 5EU law, widely copied by state and sector-specific US rulesPurpose limitation, data minimization, storage limitationFields default to "not collected" unless mapped to a lawful basis
OECD Privacy Guidelines (1980, revised 2013)International, adopted in some form by dozens of countriesCollection limitation, use limitation, accountabilityA baseline checklist for any new intake form or API
NIST Privacy Framework (2020)US, voluntary but widely adopted by federal and state agenciesRisk-based privacy controls tied to program outcomesShared vocabulary with your security and compliance partners

None of these frameworks were written with product velocity in mind, and that's the point. A Privacy Act System of Records Notice, once filed, effectively becomes a spec — it tells you what data you're allowed to hold and why, and amending it is slower than shipping a feature flag.

Much of this exposure gets locked in before a single screen gets designed, often inside the same conversations that shape how government procurement rules shape what you can build. A data-sharing clause buried in an RFP or an interagency agreement can commit your product to a retention policy months before product management is even in the room.

Privacy sits alongside a small set of other legal non-negotiables in this work, the same way Section 508 accessibility compliance functions as a legal floor, not a nice-to-have. Neither is a backlog item you deprioritize under deadline pressure — both are conditions of the product existing at all.

Purpose Limitation: The Question Every Feature Must Answer

Purpose limitation means data collected for one stated reason cannot silently migrate to serve a different one, even inside the same agency. It's the single principle most responsible for the "wait, we can't do that with this data" conversations that stall roadmaps late in a sprint.

The failure mode is rarely malicious. It's a fraud-detection field that eligibility staff start using for general case triage. It's a phone number collected for appointment reminders that ends up feeding an outbound enforcement campaign. Each reuse feels incremental; none of them were authorized by the original collection notice.

A borrowed discipline helps here more than a legal checklist does. Treat every data field like a hired employee and ask what job it's actually hired to do — the same interrogation taught in the complete guide to Jobs to Be Done. If you can't name the specific job a field performs in the specific feature you're building, you can't justify collecting or reusing it.

Run this test before any feature that touches an existing field for a new purpose:

  1. What was the original stated purpose at collection, per the notice or consent language?
  2. Is the new use a natural extension of that purpose, or a different program entirely?
  3. Does a statute, regulation, or interagency agreement explicitly authorize the new use — not "nothing prohibits it," but an affirmative basis?
  4. Would the citizen who provided this data reasonably expect this new use, given the context they gave it in (Nissenbaum's test again)?
  5. If the answer to #3 is unclear, who owns the decision — legal, privacy office, program leadership — and have you actually routed it there?

If any answer is "no" or "not sure," the feature needs a new legal basis before it needs a design review.

Retention Schedules: Data Has an Expiration Date

Retention is the most neglected half of data governance because deletion has no visible feature and no demo. Nobody celebrates a records-disposal job running successfully; the failure mode — a decade-old dataset nobody remembers the justification for — is invisible until a breach, audit, or FOIA request finds it.

Records-management authorities like the National Archives and Records Administration (NARA) publish General Records Schedules precisely because agencies otherwise default to keeping everything indefinitely. The pattern below is illustrative, not a universal rule — actual windows vary by program and jurisdiction, and your agency's records officer has the authoritative schedule.

Data CategoryTypical Legal BasisIllustrative Retention PatternDisposal Trigger
Identity verification documentsProgram eligibility statuteLife of case plus a fixed post-closure windowCase closure and the statutory window elapses
Financial or income recordsBenefits or tax administration statuteMulti-year, tied to the agency's records scheduleDetermination finalized and the appeal window closes
Communication logs (calls, chat, email)Service-delivery policyCommonly shorter than the underlying case fileCase closure or a defined service-log retention policy
Fraud investigation recordsInspector-general or law-enforcement authorityA separate, often longer schedule than the case itselfInvestigation closed and any statute of limitations passes
Aggregate or de-identified analytics dataProgram-evaluation authorityCan extend past individual case retention once properly de-identifiedNot applicable once de-identification is verified and re-identification risk is reviewed

Read across the table and one thing jumps out: retention isn't one number, it's a portfolio of different clocks running against different legal bases inside the same case. A product that treats "the case record" as a single retention unit will either delete evidence it's still legally required to hold, or keep communication logs years past their own policy window.

Legacy systems are where these mismatched clocks accumulate for decades — fields nobody can trace back to a statute, retention jobs that were never built, schedules that predate the system itself. That's exactly the terrain a public-sector legacy modernization strategy has to reckon with before any migration, because you cannot migrate a retention obligation you can't first identify.

Treat a missing retention trigger the same way you'd treat a missing acceptance criterion: the feature isn't done until deletion has an owner and a date, not just collection.

Access vs. Authorization: The Line Most PMs Miss

Technical access and legal authorization are two different questions, and conflating them is where most well-intentioned governance programs break down. A caseworker's role permissions might let them query every program a citizen has ever touched; the statute governing their specific program almost never authorizes using that full view.

This distinction rarely shows up in a data model, because access control and legal authorization live in different systems entirely — one in your identity and access management (IAM) layer, one in a policy document your engineering team has probably never read.

ScenarioCan They Access It?Are They Authorized to Use It?
Caseworker views cross-program case history via an admin query toolYes — role permits broad read accessUsually no — statute scopes use to the specific program they administer
Fraud analyst reuses a benefits application field for an unrelated enforcement leadYes — same database, same schemaNo, absent an explicit interagency data-sharing agreement
Support agent reads a citizen's full communication log to resolve one ticketYes — support tooling grants full historyPartially — only the portion relevant to the current service request
Analytics team joins two programs' datasets on a shared citizen identifierYes — identifiers match, join is technically trivialNo, unless both programs' authorizing statutes explicitly permit the linkage

Notice the pattern: the technical answer is almost always "yes," and the legal answer is almost always "it depends on a document your product team hasn't read." That gap is exactly where PMs need to build friction into the product deliberately — narrower default views, purpose-tagged queries, audit logs that capture why a record was accessed, not just that it was.

  • Design defaults for the narrowest authorized view, and require an explicit, logged reason to expand it — never the reverse.
  • Tag data access requests with a stated purpose at the point of query, so an audit can reconstruct intent later, not just activity.
  • Treat "the admin tool lets you see everything" as a bug report, not a feature — broad internal visibility is usually an accident of shared infrastructure, not a deliberate authorization decision.

The Entity-Modeling Exercise That Exposes Over-Collection

Naming every entity and relationship in a citizen dataset — out loud, in a diagram, before writing a single query — reliably surfaces over-collection that a requirements document never would. Fields hide inside vague entity names; once you have to label the entity precisely, the unjustified fields have nowhere left to hide.

Take a typical benefits-eligibility case and sketch its entities:

  • Citizen (the applicant)
  • Household (linked citizens sharing an eligibility determination)
  • Application (a single submission, tied to a household)
  • Eligibility Determination (a decision, referencing verified evidence)
  • Benefit Award (tied to a determination)
  • Payment (tied to an award)
  • Case Note (attached to an application)
  • Communication Record (tied to the citizen, across channels)
  • Fraud Flag (often, tellingly, tied to the citizen — not the case)

Two entities in that list are where over-collection hides in plain sight. The Citizen entity is where every field ever requested across every program tends to accumulate — a biometric photo captured for one pilot, an immigration-status field needed by exactly one benefit type, a veteran-status flag nobody remembers assigning ownership of. The Fraud Flag entity is worse: it's frequently modeled as citizen-scoped rather than case-scoped, which quietly turns a single program's suspicion into a cross-program blacklist nobody authorized.

The exercise works because modeling forces specificity that a feature request never does. "We need identity info" survives a planning meeting. "We need a national_id_number column on the shared Citizen table, readable by every program's caseworkers" does not — once it's written down that plainly, someone in the room asks why.

Mapping where each data point actually enters the system across a citizen's end-to-end journey usually reinforces the same finding: the same income document gets requested at three separate touchpoints because no single entity owns the authoritative copy, and each intake team quietly re-collects rather than re-using what already exists.

A Data-Minimization Decision Framework

Once the entity model is on the table, run every field on the Citizen and Household entities through this six-question test before it ships:

  1. Necessity: Is this field required to make the specific determination this feature performs — not "might be useful someday"?
  2. Least-identifying proxy: Could a category, a range, or a token substitute for the raw identifier and still satisfy #1?
  3. Purpose lineage: Which statute or program explicitly authorizes collecting this field, and does that authorization survive if the record is later joined with another program's data?
  4. Downstream exposure: Which other systems, reports, or integrations will touch this field — and does each one have its own lawful basis, rather than inheriting one from the original collection?
  5. Retention trigger: What event ends the need for this field, and is a disposal job actually scheduled against that trigger, or just assumed?
  6. Re-identification risk: Combined with the rest of the record, does this field single out an individual in a way disproportionate to the program's actual purpose?

A field that fails #1 gets cut. A field that fails #3 or #5 gets escalated to legal or the privacy office before it ships — not after an audit finds it.

Where a Data Model Makes This Governance Concrete

Everything above is easier to say than to do inside a real product, because most teams never actually draw the entity diagram — they infer it from whatever the database already looks like. Prodinja's Data Modelling studio is built around forcing that step explicitly, before a single table exists.

You name every entity, every relationship, and every field in a citizen dataset before generating schema. That's designed to surface exactly the kind of citizen-scoped fraud flag or all-programs identifier described above while it's still a diagram, not a migration.

It's a modeling discipline, not a compliance verdict — the legal judgment calls in the framework above still belong to your privacy office. But naming what you're collecting is the prerequisite for deciding what you're actually allowed to keep.

Key Takeaways

  • You're a custodian, not an owner — citizen data is held under legal mandate, not accumulated as a growth asset, because citizens can't opt out of the interaction that generated it.
  • Anchor every data decision to a named legal basis — Privacy Act of 1974-style systems-of-records notices, GDPR Article 5, OECD Privacy Guidelines, or your jurisdiction's equivalent — not to "nothing prohibits it."
  • Purpose limitation means one field, one job — if you can't name the specific authorized purpose a field serves in this feature, it isn't ready to ship.
  • Retention is a portfolio of clocks, not one date — identity documents, financial records, communication logs, and fraud files each run on different legal timelines inside the same case.
  • Access and authorization are different questions — broad technical visibility is usually an infrastructure accident, not a deliberate, legally reviewed grant of authority.
  • Entity modeling is a governance tool, not just a technical one — naming a field precisely on a diagram exposes over-collection that a requirements doc lets slide.

Frequently Asked Questions

What is the difference between data privacy and data governance in government products?

Data privacy sets citizen-facing limits on collection and use; data governance is the internal system of ownership, retention rules, and access controls that enforces those limits day to day. Privacy answers "should we collect this"; governance answers "who owns it, how long do we keep it, and who's authorized to touch it."

How does GDPR apply to US public-sector agencies?

Most US public-sector agencies aren't directly bound by GDPR unless they process EU residents' data, but its principles — purpose limitation, data minimization, storage limitation — are widely mirrored in US state privacy laws and federal guidance like the NIST Privacy Framework. Treat GDPR's structure as a well-tested reference model even where it isn't the controlling law.

What is purpose limitation in plain terms?

Purpose limitation means data collected for one stated reason cannot be reused for a materially different reason without new, explicit legal authorization. A phone number collected for appointment reminders, for example, isn't automatically fair game for an unrelated enforcement outreach campaign.

How long should a government agency keep citizen data?

There's no single answer — retention windows are set by program-specific statutes and your agency's records schedule (often issued or modeled on NARA guidance), and different data types inside the same case can carry different retention periods. Identity documents, financial records, and communication logs commonly follow different clocks even within one case file.

Who is accountable when citizen data is over-collected or misused?

Accountability typically sits with the agency as data controller under its governing statute, but product management carries real responsibility for the fields and access patterns actually built into the system. Government Accountability Office (GAO) reviews have repeatedly found agencies with gaps in privacy or records-retention controls, which is exactly the exposure a disciplined entity-modeling and purpose-limitation practice is designed to close.