A tracking plan, not a testing calendar, is the actual prerequisite for experimentation: if event names drift, properties go missing, or the same action fires twice, every test built on that data inherits errors no p-value will catch. Growth PMs who skip instrumentation design are running experiments on numbers they cannot fully trust.

Quick Answer: Before running another test, document a tracking plan (every event, its trigger, its properties, its owner), enforce one event-naming convention, and maintain a single source-of-truth spec that product, engineering, and analytics all read from. Skip this and your experiment results measure broken data, not user behavior.

Why Instrumentation Debt Silently Invalidates Every Experiment

Instrumentation debt kills experiment trust because bad tracking rarely fails loudly. A duplicated pageview, a renamed button click, or a property that quietly stopped firing after a refactor all bias results without throwing an error — so teams optimize toward noise and call it a win.

This is the uncomfortable part of growth work that gets the least attention. Roadmaps, experiment backlogs, and funnel ownership all assume the data underneath is sound. Ronny Kohavi, who ran experimentation platforms at Microsoft and Airbnb and co-authored Trustworthy Online Controlled Experiments, built his framework around a blunt premise he calls Twyman's Law: any number that looks surprisingly good is probably wrong, not a breakthrough.

His team treats a Sample Ratio Mismatch (SRM) check — verifying that traffic actually split the way the experiment configuration said it would — as a mandatory automated gate. Published research from large-scale experimentation programs has found SRM severe enough to invalidate a result showing up in a low single-digit percentage of live tests: small in isolation, but catastrophic for any specific experiment it hits, and invisible unless someone is checking for it.

Most instrumentation debt shows up in a handful of repeatable shapes:

  • Silent event loss — a redeploy drops a tracking call and nobody notices until a monthly report looks strange.
  • Duplicate firing — a client-side bug fires the same conversion event twice, inflating a variant's apparent lift.
  • Definition drift — two teams both ship a Signup Completed event, but one fires it on form submit and the other on email verification.
  • Property inconsistency — one client sends plan_type, another sends planType, and the join between them silently fails.
  • Sample ratio mismatch — the traffic split isn't what the experiment config says it should be.

None of these throw an error. They just quietly change what the dashboard says, which is exactly why they're dangerous — a broken pipeline and a genuine result look identical until someone reconciles the numbers by hand.

Symptom you noticeWhat it looks like in dashboardsLikely instrumentation cause
Winning variant, no downstream liftConversion event climbs but revenue or retention doesn't moveDuplicate or over-firing event on the variant path
Numbers that don't reconcile across toolsAnalytics tool shows one signup count, the product database shows anotherInconsistent event definitions or a missing server-side event
A metric jumps with no experiment liveFunnel-step conversion moves overnight with no code or test changeTracking code broke, or an event was silently redefined
Traffic split looks uneven52/48 instead of the intended 50/50Broken randomization, or bot traffic hitting one arm disproportionately (an SRM)

The takeaway from that table is simple: most instrumentation failures look like wins or like noise, never like failures. That's the whole argument for treating measurement as a prerequisite rather than something you patch after a test looks odd.

What a Tracking Plan Actually Is (and Why a Spreadsheet Still Works)

A tracking plan is a documented, versioned list of every event a product fires — what triggers it, what properties it carries, and who owns it. It functions as the contract between product, engineering, and analytics, and it doesn't need special software; a well-maintained spreadsheet beats an elegant tool nobody updates.

The customer data platform Segment popularized this artifact under the name "tracking plan" through its Protocols product, and the concept has outlived any single vendor because the problem it solves — three teams instrumenting the same funnel three different ways — is universal. Whether it's part of the broader growth PM role or a shared responsibility with data engineering, someone has to own a document that says, unambiguously, what "signup" means.

What Belongs in Each Row

A usable tracking plan row answers five questions at once: what is this event, when does it fire, what does it carry, who's accountable, and is it still alive. Skip any one of these and the plan degrades into a list nobody trusts enough to consult.

Event nameTriggerKey propertiesOwnerStatus
Signup StartedUser submits the signup formsource, plan_type, referrerGrowth PMActive
Signup CompletedEmail verified and account createdsignup_method, days_since_startedGrowth PMActive
Trial ActivatedUser completes first core action in trialtrial_length, activation_featureGrowth PMActive
Upgrade CompletedPayment succeeds on a paid planplan_tier, mrr_delta, discount_codeGrowth PMActive
Feature AdoptedUser uses a flagged feature 3+ times in 7 daysfeature_name, adoption_dayFeature PMActive
Checkout Abandoned (legacy)Old event, replaced by Checkout Exited——Deprecated

The Status column matters as much as the event name. A plan with no deprecation state slowly fills with events nobody trusts, which is its own form of debt — analysts start guessing which version is current instead of checking a source of truth.

Designing an Event Taxonomy That Survives Contact With Reality

An event taxonomy is the hierarchy that groups events by object and action instead of by feature, so a new feature extends an existing category instead of spawning its own one-off vocabulary. Get this structure right early and the tracking plan scales with the product; get it wrong and every new feature launch adds another dialect to the data.

Object-Action Structure

The convention most experimentation-mature teams converge on — and the one Segment's naming spec recommends — pairs a noun object (Signup, Order, Trial, Invoice) with a verb action, almost always in past tense for something that already happened (Started, Completed, Failed, Abandoned). This alone prevents the most common taxonomy failure: five different verbs for the same underlying moment because five different engineers named it independently.

Anchor Events to Moments That Matter, Not Every Click

Not every UI interaction deserves an event. The useful filter is whether an interaction represents progress on a job the user is trying to get done — the same lens behind Jobs to Be Done thinking, which asks what progress looks like rather than what a screen does.

Map candidate events against the stages in a customer journey and keep the ones that mark a real transition — activation, a value moment, a churn signal. Drop the ones that only describe UI mechanics, like a hover or a scroll depth nobody will act on.

Properties Carry Context, Not Identity

Properties describe the instance of an event, not the person doing it. plan_type, discount_code, and funnel_step belong on the event; a user's identity, traits, and lifecycle stage belong on a separate user object joined by a stable ID. Teams that stuff identity into event properties end up with the same fact duplicated across thousands of rows, and no clean way to update it retroactively when a user's plan changes.

A taxonomy that respects this split typically has three layers:

  1. Object categories — the small, stable set of nouns your product actually has (Account, Subscription, Trial, Content, Invite).
  2. Controlled verbs — a short list of actions reused across every object (Started, Completed, Failed, Viewed, Abandoned, Adopted).
  3. Shared property vocabulary — property keys defined once and reused everywhere (plan_type always means the same thing, everywhere it appears).

Naming Conventions: The Rules That Prevent 40 Versions of "Signed Up"

A naming convention is only useful if it's enforced consistently — the specific rules matter less than everyone actually following them. Pick Object Action in Title Case (or object_action in snake_case, engineering's usual preference), use past tense for completed actions, and never let versioning, dates, or environment names leak into the event name itself.

Analytics evangelist Avinash Kaushik, author of Web Analytics 2.0 and long-time digital marketing lead at Google, has spent years arguing that most organizations quietly assume their data is clean without ever checking — an assumption a strict naming convention exists specifically to remove. A convention doesn't make data trustworthy on its own, but it removes the single largest source of ambiguity: not knowing if two differently-named events mean the same thing.

Seven rules cover most of what goes wrong in practice:

  1. Pick one case style and never mix it — Order Completed everywhere, or order_completed everywhere, never both in the same plan.
  2. Object first, verb second — matches how people actually query and filter event lists.
  3. Use past tense for completed actions — Completed, not Complete or Completing.
  4. Maintain a controlled vocabulary of verbs — a fixed list (Started, Completed, Failed, Viewed, Clicked, Abandoned, Adopted) that any new event must draw from.
  5. Never encode versions, dates, or environments in the name — signup_v2 belongs in a version property, not the event name.
  6. No PII in event names or property keys — an email address or full name should never appear as a literal key or value.
  7. Deprecate, don't delete — mark an old event Deprecated and keep it firing through the migration window instead of breaking historical queries overnight.
Anti-patternWhy it breaksConvention-compliant version
btn_click_signup_v2Encodes UI mechanics and a version number instead of the underlying momentSignup Started with {source: "homepage_button", version: "v2"}
SignUpCompleteMixed case, present tense, no space — a query-matching nightmareSignup Completed
user_upgraded_to_pro_jan2025Bakes a date and a plan name into the event itselfPlan Upgraded with {plan: "pro", upgraded_at: <timestamp>}
clickTells you nothing about the object or outcomeFeature Adopted with {feature: "csv_export"}

This is also, quietly, one of the differences between a growth PM and a core PM: growth work lives or dies on whether the underlying event stream is trustworthy at scale, so naming discipline isn't a nice-to-have, it's part of the job description.

The Case for a Single Source-of-Truth Tracking Spec

A single source-of-truth spec matters because the moment two teams maintain separate versions of "what events exist," they silently diverge — and every dashboard, experiment readout, and model built on top inherits that disagreement without anyone noticing until the numbers stop reconciling.

A tracking plan that lives in three different documents isn't really a tracking plan. It's three guesses, and only one of them can be right.

The practical fix looks a lot like how engineering teams already manage schemas: version the spec, review changes before they ship, and treat "add a new event" as a change request, not a Slack message. Teams that get experiment velocity right without drowning in QA overhead almost always have this governance layer in place — it's what lets testing throughput scale without every new test requiring a fresh manual audit of what's actually being tracked.

A workable governance loop has four steps:

  1. Propose — anyone adding an event drafts the row (name, trigger, properties, owner) before writing code.
  2. Review — a designated owner (often the growth PM) checks it against the existing taxonomy and naming convention.
  3. Ship — engineering implements against the approved spec, not an ad hoc Slack thread.
  4. Audit — periodically diff what's actually firing in production against what the spec says should be firing.

Where a Data Modelling Tool Fits

The gap between "we agreed on a clean event schema in a meeting" and "the schema that actually got built" is where most instrumentation debt is born. The spec lives in a doc, the implementation lives in code, and nobody keeps them in sync by hand for long.

This is the specific handoff Prodinja's Data Modelling tool is built for: you define your event entities and properties, and it turns that model into SQL DDL you can hand straight to a data engineer.

The idea is to shorten the distance between the schema you designed and the schema that gets built, instead of leaving that translation to six people's slightly different interpretations. It's meant as a way to get the tracking schema right before the analytics get messy, not a fix for a taxonomy that's already drifted.

Auditing and Rolling Out Instrumentation Without Stalling the Roadmap

An instrumentation audit is a bounded, periodic review, not a full rebuild. The fastest path to a trustworthy baseline is inventorying what's actually firing in production today, tagging each event active, duplicate, or dead, then fixing the highest-leverage funnel events first instead of trying to perfect everything at once.

A practical sequence for a first audit:

  1. Pull every event actually firing in production over the last 30 days — not the tracking plan, reality.
  2. Diff that list against the documented tracking plan and flag anything undocumented, duplicated, or missing.
  3. Rank the gaps by funnel criticality — fix acquisition, activation, and revenue events before anything cosmetic.
  4. Add an instrumentation check to the definition of done for new features, so debt stops accumulating while you clean up the backlog.
  5. Set a recurring audit cadence (quarterly is common) with a named owner, usually the growth PM closest to the funnel.

Rolling this out doesn't require a moratorium on shipping. Fix the events feeding your highest-priority experiments first, deprecate the worst offenders on a visible timeline, and treat the rest as backlog — a tracking plan that's 80% clean and actively maintained beats a perfect one that took two quarters and blocked every other launch.

Key Takeaways

  • Instrumentation debt fails silently — broken tracking produces plausible-looking numbers, not errors, so it invalidates experiments without anyone noticing until results stop reconciling.
  • A tracking plan is a contract, not a wishlist — document every event's trigger, properties, owner, and status, and a spreadsheet is a perfectly good place to start.
  • Object-Action naming prevents taxonomy sprawl — pair a stable noun with a controlled, past-tense verb so five teams don't invent five names for the same moment.
  • Properties hold context, not identity — keep user traits on a stable user object and event-specific facts on the event itself.
  • Governance beats heroics — a versioned, reviewed source-of-truth spec is what lets experiment velocity scale without a manual audit before every test.
  • Audits should be bounded and recurring, not a one-time rebuild — inventory reality, diff against the plan, fix the highest-leverage gaps first.
  • Measurement comes before testing, not after — a growth PM's first job on a new surface is making sure the events are trustworthy, then running the experiment.

Frequently Asked Questions

What is a tracking plan in analytics?

A tracking plan is a documented, versioned specification of every event a product fires — including its trigger, its properties, and its owner — that product, engineering, and analytics all treat as the single reference for what's being measured. It can live in a spreadsheet or a dedicated tool; what matters is that it's actually kept current.

How do you name analytics events consistently?

Pick one case style (Title Case or snake_case), structure names as Object Action with a controlled, past-tense verb list, and keep versions, dates, and environment details in properties instead of the event name. Consistency matters more than which specific style you pick, since the goal is removing ambiguity about whether two events mean the same thing.

What is instrumentation debt, and why does it matter for experimentation?

Instrumentation debt is the accumulated gap between what a tracking plan says a product measures and what's actually firing correctly in production. It matters because experiments run on top of that gap inherit its errors invisibly — a test can show a confident, statistically significant winner that's actually an artifact of duplicate or missing events.

Who should own the tracking plan — product, engineering, or analytics?

In most functioning setups, a growth PM owns the tracking plan as a product artifact while engineering implements against it and analytics validates the output, similar to how a spec or PRD gets owned. Shared ownership without a single accountable owner is usually how tracking plans drift in the first place.

How often should a team audit its event tracking?

A quarterly cadence is a reasonable default for most product teams, with an additional lightweight check whenever a major redesign or funnel change ships, since refactors are the single most common point where events silently break or get renamed. Teams running frequent experiments may want a tighter, monthly cadence on the events feeding active tests.