Great engagement metrics can coexist with zero learning gain, because streaks, session length, and daily actives measure whether someone opened the app — not whether they know more than they did last week. The fix: track two separate columns, fast engagement leading indicators and slower learning lagging indicators, and check them against each other every review cycle, not just the first one.
Quick answer: Streaks,
DAU, and time-on-app are leading indicators that predict learning only until users learn to optimize them directly. Pair them with lagging indicators —pre/post assessment delta, 30-day retention of concept, and transfer to novel problems — and watch the gap between the two columns, not either column alone.
Why Engagement Is a Proxy Metric That Decays Under Optimization Pressure
Engagement metrics decay as learning signals precisely because they work — the moment you reward a behavior, users find the cheapest path to producing it. That's Goodhart's Law: economist Charles Goodhart observed that once a measure becomes a target, it stops being a reliable measure, a point anthropologist Marilyn Strathern later sharpened into the version most product teams quote today.
Streaks are a textbook target. They're visible on the home screen, gamified by design, and the easiest number on your dashboard to move without touching the product's actual value.
- Streaks and
DAUreward opening the app, not thinking hard inside it. - Session length rewards staying, even if the extra minutes are spent re-reviewing content already mastered.
- Lessons completed per day rewards volume, which is trivially inflated by choosing easier content.
None of these are bad metrics to watch. They're bad metrics to manage toward, because the moment a roadmap optimizes for them directly, users find the shortest path to producing the number without producing the behavior it was supposed to represent.
The Reinforcing Loop Nobody Diagrams
Here's the mechanism, stated as a loop rather than a list. You notice streaks correlate with retention, so you invest in streak mechanics — reminders, freezes, badges — and streaks climb. The team celebrates, and the roadmap funds more of the same mechanic next quarter.
But some learners now maintain their streak by doing the least effortful thing that counts: reviewing content they've already mastered instead of tackling new material. Learning quietly decouples from the streak, and nobody notices right away, because the number everyone is watching is still climbing.
That's a reinforcing loop: the more you optimize the proxy, the more disconnected it becomes from the thing it was meant to represent — and celebrating the proxy funds even more investment in it.
Duolingo's own growth team has written publicly about this exact tension, adding features like streak freezes and streak repair specifically because rigid streak pressure was pushing some users toward anxiety-driven, low-value app opens rather than genuine practice. A large public product acknowledging the fragility of its own headline metric is a useful signal that this isn't a hypothetical risk.
Why This Bites AI Tutors Harder Than Static Courseware
An AI tutor doesn't just get gamed by users — its own personalization layer can drift toward the same proxy if nobody's watching. If a recommendation or difficulty-adjustment model is tuned on engagement signals like session length or completion rate, it will learn to serve whatever keeps sessions long, which is not reliably the same content that produces learning gains.
That's Goodhart's Law operating twice in the same system: once in the human choosing the easiest path, and once in the algorithm optimizing toward the same easy path because that's what its training signal rewards. A pre/post assessment delta that's flat while a recommendation model's engagement score climbs is worth investigating as a modeling problem, not just a user-behavior problem.
This Isn't Only an EdTech Problem
The same Goodhart trap shows up wherever a product substitutes a cheap, fast signal for an expensive, slow truth:
- Fintech apps optimizing "sessions per week" over actual financial health, a pattern our fintech product management guide covers in more depth.
- Healthtech apps rewarding medication-reminder taps instead of adherence outcomes, discussed in our healthtech product guide.
- Ecommerce teams chasing session count over repeat-purchase value, per our ecommerce and retail guide.
- Martech and adtech dashboards celebrating click-through rate while conversion quietly erodes, covered in our martech and adtech guide.
If you're building the broader edtech roadmap around this problem, our complete guide to edtech product management is the fuller playbook this article zooms into.
The Two-Column Metric Tree: Engagement Leading Indicators vs Learning Lagging Indicators
A metric tree fixes the Goodhart trap by refusing to let one number stand in for the whole story. Build two columns — engagement metrics that move daily and predict interest, and learning metrics that move slowly and prove mastery — then require both to move together before you call anything a win.
| Engagement leading indicators (fast, cheap, gameable) | Learning lagging indicators (slow, expensive, harder to fake) |
|---|---|
DAU / WAU and streak length | Pre/post assessment delta (mastery score change) |
| Session length / time-on-app | 30-day retention of concept (delayed post-test) |
| Lessons or problems completed per day | Transfer to novel problems (near/far transfer accuracy) |
| Push notification open rate | Time-to-mastery (attempts to reach criterion) |
| Streak-save / streak-freeze usage | Error-pattern shift (specific misconceptions resolved) |
Benjamin Bloom's landmark 1984 "Two Sigma Problem" found that one-on-one tutoring could move average students roughly two standard deviations above students in conventional instruction. That's the scale of gain a serious learning lagging indicator is measuring against — not a vanity target, a genuinely high bar.
Defining Each Learning Lagging Indicator
Each right-column metric needs an operational definition or it turns into a vague aspiration:
Pre/post assessment delta— score on a held-out assessment before a unit versus after, using items the tutor never explicitly drilled.- 30-day retention — accuracy on the same concept re-tested a month later, without warning and without hint assistance.
- Transfer to novel problems — accuracy on problems that apply the same underlying concept in an unfamiliar format or context.
These are all harder to collect than a DAU count, and that's the point. Psychologist Robert Bjork's research on "desirable difficulties" is exactly why: conditions that feel effortful and slower tend to produce durable learning, while conditions that feel smooth and fast often just look good on a dashboard.
Reading the Two Columns Together
The tree only earns its keep if you read both columns as one chart, not two separate dashboards:
- If both columns rise together, the product is working as intended.
- If engagement rises while learning flatlines, you're watching Goodhart's Law happen in real time.
- If learning rises while engagement dips, you may have added rigor that's correctly filtering out low-effort sessions.
A Worked Example: How a Learner Games the Tutor Without Breaking a Single Rule
Say a learner is assigned daily practice in an AI tutor that mixes new material with review of previously mastered items, and the streak counts either kind of activity equally. Within a few weeks, that learner discovers the review queue is faster and easier than the new-material queue.
They start opening the app, clearing a few minutes of easy review items, and closing it — never touching harder new content, never risking the streak on a problem they might get wrong.
What the Dashboard Shows vs What Actually Happened
On the growth dashboard, this learner looks great:
- Streak: 47 days and climbing.
DAU: counted, every single day, on schedule.- Lessons completed: at or above the cohort median.
- Accuracy: high, because review items are, by definition, already mastered.
On a delayed post-test of the unit's actual content, this same learner's pre/post assessment delta is close to flat. Nothing new got learned in a meaningful stretch of those seven weeks — the tutor became a low-stakes habit loop, not a study tool.
A Second, Quieter Way to Game a Tutor
A subtler version of the same behavior: a learner who taps every hint in sequence until the tutor reveals the answer, then submits it. Researchers Ryan Baker, Albert Corbett, and Kenneth Koedinger documented exactly this pattern in classic studies of intelligent tutoring systems, where students who "gamed" hint sequences to reach correct answers without reasoning learned reliably less than peers solving the same problems honestly — even though both groups looked identical on completion metrics.
Neither behavior breaks a single product rule. Both quietly sever the link between the metric being celebrated and the outcome the product actually exists to produce.
The Guardrail Metrics That Catch Gaming Before It Wrecks Your Roadmap
A guardrail metric doesn't replace an engagement metric — it sits next to it and flags when the composition of the behavior producing that number has changed. You're not trying to catch individual learners; you're trying to catch a cohort-level shift in how the number gets produced.
| Gaming behavior | What engagement metrics show | Guardrail metric that catches it |
|---|---|---|
| Streak-farming via easy review only | Streak length and DAU stay strong | New-to-review item ratio per session drops |
| Hint bottom-outing | Completion rate and accuracy look strong | Bottom-out hint rate (all hints used before answering) rises |
| Fast guessing on multiple choice | Items completed per session rises | Median response time per item falls below a plausible reading floor |
| Streak-saving cramming bursts | Streak stays unbroken | Streak-save timing clusters in the final hour before reset |
Building a Guardrail Metric in Four Steps
- Define the expected behavior distribution — what does honest, effortful use of this feature actually look like, in numbers, for a typical cohort?
- Name the implausible signature — a response time too fast for genuine reading, a hint sequence too complete to be anything but a shortcut.
- Set the threshold as a rate, not a raw count, so it scales with usage instead of triggering false alarms as your user base grows.
- Review at the cohort level, not the individual level. The goal is catching a systemic incentive problem, not surveilling single learners — which also keeps you clear of
FERPAandCOPPAobligations around monitoring minors.
Keep the guardrail's exact thresholds internal rather than surfacing them in-product. A guardrail metric that learners can see and reverse-engineer becomes a second target to game — the same Goodhart problem one level down, just with a smaller, more sophisticated audience.
Guardrail metrics are cheap insurance relative to what they catch. A single new-to-review ratio, tracked weekly, would have caught the streak-farming example above months before a quarterly business review had to explain a flat pre/post assessment delta sitting next to a record streak chart.
Operationalizing the Metric Tree: Ownership, Cadence, and Seeing the Loop
The metric tree fails if it lives in a slide nobody revisits. Give each column an owner and a cadence: growth or lifecycle PMs own the engagement column and review it weekly, while core-product or learning-science leads own the lagging column and review it on the natural rhythm of your assessment design — typically every two to six weeks.
Neither owner should present their column without the other one on the same page. A weekly growth readout that never shows the pre/post assessment delta trend line is exactly how a reinforcing loop goes unnoticed for two quarters.
This also has implications for how you write OKRs. Make the learning lagging indicator the committed objective, and treat the engagement leading indicator as a supporting health metric with a floor, not a target with a ceiling to chase. That single ordering choice is often what keeps a roadmap review from quietly drifting back toward optimizing the proxy.
Where Systems Thinking Earns Its Keep
The hardest part of this problem isn't picking metrics — it's seeing the loop connecting them before it does damage. Mapping the system explicitly, rather than describing it in a paragraph on a slide, tends to change the conversation in a roadmap review.
Optimize streak → learners game streak → learning decouples → team celebrates the streak anyway. Diagrammed as a loop, that structure reads as reinforcing — the kind that gets worse, not better, the longer it goes undiagrammed.
You build the diagram; the canvas does the reading, flagging the loop type from the polarity and direction of the links you've drawn rather than analyzing your live product data on its own. That's often enough to get a guardrail metric funded before the next roadmap cycle, not after one already looks bad.
If your tutoring product also serves parents or school admins who see a different slice of this same data, our guide to designing for learner, teacher, admin, and parent covers how to keep each audience's dashboard honest without contradicting the others.
Key Takeaways
- Engagement metrics are leading indicators, not proof of learning — streaks,
DAU, and session length predict interest, not mastery. Goodhart's Lawis the mechanism, not a metaphor: once a proxy becomes the target, users find the cheapest path to producing it.- Build a two-column metric tree: engagement leading indicators paired with learning lagging indicators like
pre/post assessment delta, 30-day retention, and transfer to novel problems. - Gaming rarely looks like cheating — it looks like streak-farming easy review items or bottoming out hints, both invisible on a standard engagement dashboard.
- Guardrail metrics catch composition shifts, not individual bad actors — track rates like new-to-review ratio and bottom-out hint rate at the cohort level.
- Diagram the reinforcing loop before it compounds. Seeing "optimize streak → gaming → decoupled learning" as a loop makes the case for a guardrail metric far faster than a paragraph does.
Frequently Asked Questions
Is engagement always a bad signal to track for an AI tutor?
No — engagement is a legitimate, useful leading indicator; it's only a problem when it's the only thing on the dashboard or the direct target of a roadmap. Pair every engagement metric with a learning lagging indicator and require both to move together before calling a feature a win.
How often should you run pre/post assessments in an AI tutoring product?
Run a pre/post assessment around every content unit or module boundary, and add a 30-day delayed retest for concepts you consider core. Most tutoring products land somewhere between every two and six weeks, matched to how the curriculum is chunked rather than to sprint cadence.
What is Goodhart's Law and why does it apply to edtech metrics?
Goodhart's Law, from economist Charles Goodhart, says a measure that becomes a target stops being a good measure, because people optimize the measure rather than the underlying behavior. In edtech, streaks and DAU are especially exposed because gamified UX makes the proxy trivially easy to move on its own.
How do you detect students gaming an AI tutor?
Look for composition shifts rather than individual outliers: a rising bottom-out hint rate, a falling new-to-review item ratio, or response times faster than a plausible reading floor. Classic research on intelligent tutoring systems by Ryan Baker and colleagues found these patterns reliably predicted lower learning gains even when completion metrics looked normal.
Should we remove streaks and gamification from an AI tutoring product?
Not necessarily — streaks can genuinely drive habit formation, which matters for a practice product. The fix isn't removing engagement mechanics; it's refusing to let them stand alone, and adding guardrail metrics that catch the specific ways learners might farm a streak without engaging with new material. Treat gamification as a retention lever to keep instrumented, not a proof point to retire once it exists.