Survivorship bias in your cohort curves
By Isaac Turner, Measurement Editor — Archive date: 6 min read
View author profile
Your LTV curves are built from the campaigns you kept, the users who consented and the cohorts old enough to matter. Each filter inflates the tail.
The retention and revenue curves a UA team extrapolates from aren't a sample of the players it acquired. They're a sample of the players it acquired and then kept paying for, and could observe, and had time to watch; each of those conditions removes cohorts from the dataset, and none of them does so at random. The result is a curve library that describes the best-case version of the game, and every forecast built on it inherits the optimism.
This is a stronger claim than "data has noise". It says we know the direction of the error in advance. Curves built from surviving cohorts will overstate the long tail, and the overstatement is largest exactly when the team is about to do something new: open a channel, enter a geo, double the daily budget.
Four filters, all pointing the same way
The first filter is campaign pruning. A buyer kills underperforming campaigns early, usually on a D3 or a D7 signal. The killed campaigns produce small, young cohorts that never reach D90, so they contribute nothing to the D90 curve. The campaigns that survive to maturity are, by construction, the ones that looked good early, so the mature curve is a curve of winners.
The second filter is maturity truncation. Only cohorts acquired months ago have long-horizon data, and those cohorts were bought under a different budget, a different creative set, possibly a different channel mix, and often at a smaller scale where the network was still serving its best inventory. The curve's tail is always older than its head, and older usually means cheaper and better targeted.
The third filter is consent and attribution. On iOS, the users you can follow at the individual level are the ones who opted in, and opted-in users aren't a random draw from the install base. On both platforms, attributed users are the ones the rule assigned to a paid source, and heavily engaged players are more likely to have multiple touches and therefore more likely to be attributed. Unattributed installs, which may be weaker, fall into organic; they vanish from the paid curve.
The fourth filter is the model itself. Predicted LTV systems learn from cohorts that reached maturity, which means they learn from survivors of the first three filters. The model learns that early behaviour predicts a generous tail because, in the training data, it did. It then applies that multiplier to new cohorts that were never subjected to the same selection.
The mechanism that makes it compound
These filters would matter less if they were independent, and they aren't. Buyers make pruning decisions on early signals, the pLTV model converts early signals into forecasts, and the forecasts set the D7 thresholds that drive the next round of pruning. As argued in an earlier UA Ledger piece, D7 ROAS targets punish long-retention games, and here is the companion problem: the multiplier used to set that target was itself estimated from cohorts that had already passed a version of it.
The compounding shows up as a specific failure pattern. A team scales into a channel on the strength of a strong predicted D180. The realised curve on the scaled cohorts falls below the forecast from about D30 onward. The team concludes the channel degraded at scale, which is partly true, but the larger effect is that the forecast was never a forecast for this kind of cohort. It was a forecast for the cohorts that survived long enough to measure.
The second-order consequence most coverage misses is that survivorship bias makes new channels look worse than they are, relative to old ones. The incumbent channel's curve comes from years of pruned, matured, consented survivors; the new channel's comes from everything it delivered in its first six weeks and none of it pruned or mature. Comparing the two side by side is comparing a curated highlight reel to raw footage, and the incumbent wins every time. Channel diversification stalls not because the new channels are weak but because the comparison itself tilts against them.
A framework: two curve libraries
The corrective is procedural rather than statistical, and it begins with keeping what the team currently throws away.
Maintain two curve libraries. The first is the live-campaign library, the one everyone already has: mature cohorts from campaigns the team kept. The second is the all-spend library, which includes every cohort that received budget, including campaigns killed on D3, geos exited after a month and creatives paused after a week. Killed cohorts stop accruing new data, but the data they did produce belongs in the denominator.
The gap between the two libraries at any given day is a direct estimate of how much pruning has inflated the curve. An illustrative gap of five percentage points at D30 tells you that a forecast built from the live library will run roughly that optimistic for any spend nobody has pruned yet, which is all new spend.
From that gap, a decision rule follows. When forecasting for existing, pruned channels at current budget, use the live library. When forecasting for any expansion, use the all-spend library, or the live library shrunk by the observed gap. If the two libraries have diverged by more than a threshold the team sets in advance, treat extrapolation past D30 as unreliable and cap expansion spend until fresh cohorts mature.
Four further habits close the remaining filters:
- Record the acquisition context of every mature cohort: daily budget, channel mix, creative generation.
- When the head of the curve is being bought under different conditions from the tail, say so in the forecast.
- Build the iOS curve on SKAN or AdAttributionKit aggregate postbacks as well as on consented users, and compare. Where the consented curve sits well above the aggregate curve, the consented sample is not representative.
- Ask the pLTV vendor or internal team whether killed campaigns are in the training set. If not, the model is learning from winners; treat its multiplier as a ceiling.
What the numbers will do when you fix this
Expect the corrected curves to look worse. The all-spend library will show a lower tail, the forecast for new channels will come down, and the payback horizon for expansion will lengthen. That isn't the measurement getting worse; it's the measurement catching up with what the game has been doing all along, which the survivor-only view was hiding. The growth lead who accepts a lower forecast now is buying a smaller gap between forecast and outcome at D180, and that gap is the thing finance actually remembers.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·1 min read
Web-shop attributed revenue in MMP vs payment-provider settlements
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read