A churn early-warning system doesn't require machine learning to be useful — it requires catching declining engagement in leading indicators (session frequency, key feature usage, support friction) early enough to intervene. Start with simple threshold rules on 3-5 signals, combine them into a weighted health score, and pair every alert with a defined intervention. Complexity should come later, if at all.
Quick Answer: Track 3-5 leading indicators (session decline, dropped key-feature use, support tickets, invoice/payment friction), combine them into a weighted health score with tiered thresholds, and route each tier to a specific, pre-built intervention. Skip ML until rules plateau.
Most growth PMs hear "churn prediction" and assume they need a data scientist, a feature store, and a logistic regression model. That assumption stops good work before it starts. A threshold-based system built in a spreadsheet or a BI tool this week will catch most of the same at-risk users a first-pass ML model would catch six months from now — because both are built on the same underlying insight: churn is usually preceded by a visible decay in behavior, not a sudden cliff.
This piece walks through building that system: which signals to trust, how to combine them into one score, where thresholds should live, and — the part most churn write-ups skip — what happens the moment the alert fires.
Why Leading Indicators Beat Lagging Ones
A leading indicator changes before a user decides to leave; a lagging indicator confirms the decision after it's already made. Cancellation requests, downgrade clicks, and support tickets that say "how do I export my data" are lagging — useful for diagnosis, useless for early warning. The system you want watches the weeks before that moment, not the day of it.
The academic grounding here is old and solid. Fred Reichheld's work on the loyalty economics behind Net Promoter Score (Bain & Company, early 2000s) established that behavioral disengagement predates stated dissatisfaction — customers stop doing the thing that made the product valuable well before they say anything about it. Subscription-economy research popularized by Gainsight and the broader customer-success industry has repeatedly found the same pattern across B2B SaaS cohorts: usage decay is the dominant early signal, ahead of NPS score drops or support volume.
Practically, that means you're looking for decay in the specific action tied to value, not overall app opens. A user who logs in daily but has stopped touching the feature that solves their actual job is often more at-risk than one who logs in twice a week and still uses it. This is the same distinction covered in depth in the complete guide to growth and retention — leading indicators are retention's forward-looking half.
The Three Signal Families Worth Tracking
Not every metric deserves a place in your churn model. Restrict yourself to three families:
- Engagement decay — session frequency, session length, or days-active trending down over a rolling window (commonly 14-28 days) compared to the user's own baseline, not a global average.
- Key-feature abandonment — a drop in usage of the specific feature tied to your activation metric or to the aha moment you've already identified for the product.
- Friction signals — rising support ticket volume, failed payments, seat reductions, or repeated errors on a core workflow.
Threshold Rules Before Machine Learning
Threshold rules — "flag any user whose weekly key-feature usage drops more than 40% versus their own trailing 4-week average" — catch the majority of at-risk accounts with none of ML's data, tooling, or maintenance overhead. They're transparent, debuggable, and something a PM can build and adjust alone. Reach for ML only once rules plateau and you have enough labeled churn history to validate a model against.
The case for rules-first isn't just pragmatism; it's sequencing. Ron Kohavi and the broader experimentation research community have long emphasized that simple, interpretable baselines should exist before any complex model, because a baseline tells you whether the complex model is actually adding lift or just adding noise. If your threshold rule already catches 70% of churners two weeks out, an ML model's job is narrowing the remaining 30% — not replacing the whole system.
What a Threshold Rule Actually Looks Like
A workable rule has four parts: the metric, the comparison window, the trigger condition, and the user segment it applies to. For example:
| Rule component | Example value |
|---|---|
| Metric | Weekly logins |
| Comparison window | Trailing 4 weeks vs. prior 4 weeks |
| Trigger condition | ≥ 40% decline |
| Segment | Paid accounts active 60+ days |
Keep each rule segment-aware. A brand-new trial user with two logins in week one and zero in week two isn't "declining 100%" in any meaningful sense — they never established a baseline. Apply decay rules only to users who've had time to form a habit, which is also why time-to-value work matters upstream of churn work: you can't detect a drop-off from a baseline that was never set.
Where Rules Break Down
Threshold rules struggle with three things: seasonal or cyclical usage (a B2B tool used only at month-end), multi-signal interactions (declining logins plus rising support tickets is worse than either alone), and thresholds that were tuned once and never revisited. None of these require ML to fix — they require a composite score and a maintenance habit, both covered next.
Building a Health-Score Composite
A health score combines several weighted signals into one number so a single declining metric doesn't trigger noisy alerts, while multiple simultaneously weakening signals surface clearly. Weight each component by how strongly it correlates with past churn in your own data, not by gut feel, and review those weights quarterly.
A simple composite might look like this:
| Signal | Weight | Scoring logic |
|---|---|---|
| Key-feature usage trend | 35% | 0 (no decline) to 100 (feature abandoned) |
| Session frequency trend | 25% | 0 to 100 based on % decline vs. baseline |
| Support ticket volume | 20% | 0 (none) to 100 (3+ tickets in 30 days) |
| Payment/billing friction | 15% | 0 (clean) to 100 (failed payment or downgrade) |
| Seat/license change | 5% | 0 (stable/growing) to 100 (reduced) |
Multiply each component's 0-100 score by its weight and sum them for a composite risk score. Score above 60, escalate to a proactive outreach tier; 40-60, monitor closely; below 40, no action. These exact cutoffs are a starting point — recalibrate them against your own churned-cohort history within the first quarter of running the system.
Two design choices matter more than the exact weights:
- Weight by evidence, not intuition. Pull your last 6-12 months of churned accounts and check which signals actually moved before they left. If support tickets barely correlated with churn in your data, don't give them 20% weight because a churn article said to.
- Recency-weight within each signal. A decline that happened 3 days ago should count more than one from 25 days ago inside the same rolling window — a simple exponential decay or even a straight recency multiplier avoids the score feeling stale.
Composite scores also let you connect churn risk back to the customer journey — a risk spike right after onboarding reads differently, and needs a different playbook, than one six months into a mature account.
The Intervention Playbook Is the Point
A churn score with no attached action is a dashboard, not an early-warning system — the entire value of prediction is the intervention it triggers. For every risk tier, define who gets notified, what they do, and within what SLA, before you turn the scoring on. Otherwise the alert becomes noise a CSM learns to ignore within a month.
Build the playbook as a simple decision table, one row per tier:
| Risk tier | Score range | Owner | Action | SLA |
|---|---|---|---|---|
| Critical | 80-100 | CSM / Account exec | Personal outreach call, offer training or config review | 24 hours |
| High | 60-79 | CSM | Automated check-in email + in-app nudge to dropped feature | 3 days |
| Watch | 40-59 | Lifecycle/marketing | In-app tip or email drip re-surfacing the abandoned workflow | 1 week |
| Healthy | 0-39 | None | No action; keep monitoring | N/A |
A few rules make playbooks actually get used instead of ignored:
- Match the intervention to the likely cause, not a generic "check in" email. A user who dropped a specific feature needs a nudge back to that feature, not a survey.
- Cap volume per owner. If your CSM team gets 200 "critical" alerts a week, the tier boundary is miscalibrated — narrow it until the volume matches realistic outreach capacity.
- Track intervention outcomes, not just alert counts. If personal outreach at the "critical" tier isn't measurably changing retention, the tier's action needs to change, not just its threshold.
- Revisit playbooks when the product changes. A new onboarding flow or pricing tier can shift what "declining usage" even means for a cohort.
Documenting the Signals You Actually Observe
Most teams lose the compounding value of this system because the signals a CSM or PM notices qualitatively — "this account has been asking oddly basic questions for a month," "their champion just left" — never make it into the quantitative model at all. This is where Prodinja's Journals are designed to help: they let you log the risk signals you observe in real conversations and account reviews as you notice them, so a pattern of pre-churn behavior — the qualitative half of churn detection — becomes a documented, reviewable list rather than something that lived only in one CSM's memory. Over time, that log is exactly the kind of evidence you'd use to check whether a rule's weights still hold.
Common Mistakes That Undermine Early-Warning Systems
Most churn early-warning failures trace back to four causes: comparing users to a global average instead of their own baseline, alerting without an owner, never revisiting thresholds, and treating the score as precise when it's directional. Each is cheap to avoid once you know to check for it.
- Global baselines instead of individual ones. A power user's "40% decline" and a light user's "40% decline" mean different things in absolute terms. Compare each user to their own trailing history.
- No named owner per alert tier. An alert nobody is accountable for acting on trains the org to ignore the whole system within a quarter.
- Static thresholds. Product changes, seasonality, and pricing changes all shift what "normal" usage looks like; a threshold set at launch is stale within two quarters.
- Overtrusting the score's precision. A composite score of 63 isn't meaningfully different from 58 — treat scores as tiers, not rankings, and don't over-engineer decimal precision into a system built on directional signals.
- Ignoring the qualitative signals. Quantitative decay metrics miss champion departures, competitive evaluations, and stated frustration that a CSM hears directly — these deserve a place in the system too, not just a side conversation.
Key Takeaways
- Start with threshold rules on 3-5 leading indicators — engagement decay, key-feature abandonment, and friction signals — before considering machine learning.
- Compare each user to their own baseline, not a global average, and only apply decay rules once a user has had time to establish a habit.
- Combine signals into a weighted health score so isolated noise doesn't trigger false alarms while genuine multi-signal decline stands out clearly.
- Weight components using your own churned-cohort history, not assumed importance, and recalibrate thresholds at least quarterly.
- A prediction without a paired intervention is just a dashboard — define an owner, action, and SLA for every risk tier before turning scoring on.
- Log qualitative risk signals alongside quantitative ones so patterns a CSM notices in conversation don't disappear from institutional memory.
- Track intervention outcomes, not alert volume, and adjust the playbook — not just the threshold — when an action isn't moving retention.
Frequently Asked Questions
Do I need machine learning to predict churn?
No — most teams can catch the majority of at-risk users with simple threshold rules on 3-5 leading indicators before ever training a model. ML earns its complexity only once rules plateau and you have enough labeled churn history to validate real lift over the simpler baseline.
What's the difference between a churn score and a health score?
A churn score typically estimates probability of cancellation directly; a health score is broader, combining engagement, feature usage, and friction signals into one composite that reflects overall account wellness. In practice, a well-built health score functions as your churn early-warning signal without needing a separate predictive model.
How many signals should a churn early-warning system track?
Three to five is a practical range — enough to capture engagement decay, key-feature abandonment, and friction without the model becoming unmaintainable. More signals rarely improve accuracy meaningfully and make the system harder to debug when a false alarm fires.
How often should churn thresholds be recalibrated?
Quarterly is a reasonable default, with an additional review any time the product, pricing, or onboarding flow changes materially. Thresholds set once at launch drift out of relevance as "normal" usage patterns shift with the product itself.
What should happen immediately after a churn alert fires?
The alert should route to a named owner with a pre-defined action and SLA — a personal outreach call for the highest tier, an automated nudge for lower tiers. A system without this step only produces a dashboard nobody acts on, which is the single most common reason churn-prediction efforts get abandoned within a year.