Weight cohort retention by installs rather than by row count
Analysis
By Isaac Turner, Measurement Editor — 2 min read
View author profile
A spreadsheet that averages campaign percentages gives a small test and a large campaign the same influence. A reproducible synthetic example with editable inputs and explicit limits.
A spreadsheet that averages campaign percentages gives a small test and a large campaign the same influence. This lab compares that row average with the pooled retention rate. The arithmetic is simple; the decision is whether the populations are comparable enough to pool at all.
Define the calculation before using it
For each cohort, retain its eligible install count and its returning-player count. Add numerators and denominators before calculating the combined percentage. Do not average displayed percentages after they have been rounded. Check that both rows use the same retention window, identity rule and activity definition, otherwise a weighted result can be precise while answering no coherent question.
Work through the synthetic example
The synthetic small cohort returns 30 of 100 players; the large cohort returns 90 of 900. The unweighted average is 20%, while the pooled result is 12%. The worksheet displays both to reveal the discrepancy. Neither result proves that a channel improved: a changed campaign mix can change the combined figure even when each cohort’s own rate stays constant.
| Output | Worked-example result |
|---|---|
| Pooled retention % | 12.0000 |
| Unweighted average % | 20.0000 |
Use the artifact and preserve its assumptions
Open the editable calculator to change the inputs and inspect the sensitivity view. The CSV records synthetic inputs and expected outputs; the JSON fixture keeps the equations available for reproduction. These calculations have been checked against the stated example. No measured campaign data is included.
The sensitivity rows vary only n2 by 20% below and above the entered value. They are scenarios, not confidence limits or a forecast distribution. A row outside the model’s constraints is labelled rather than turned into a plausible-looking result. Save the chosen inputs with the decision so another reader can distinguish a changed assumption from a changed formula.
Evidence and limits
Google’s export schema separates the reporting date, UTC timestamp and event parameters. The worksheet uses simplified aggregate inputs; it does not claim that an export already contains a correctly reconciled cohort. See Google Analytics BigQuery Export schema, especially “event_date, event_timestamp, event_value_in_usd and event_params fields”.
Pooling is not adjustment for geography, device, audience or creative. Keep those dimensions available if the aggregate will be used to allocate spend between unlike populations.
Background: UA metrics explained: CPI, ROAS, LTV and payback and Reading an MMP dashboard without fooling yourself. These existing articles provide context; the present calculation does not verify every archived claim.
Featured
Related posts

measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team

measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios

measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)

measurement
media buying
·1 min read
View-through attribution windows on F2P rewarded and interstitial traffic
More from the Measurement desk

measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel

measurement
·2 min read
When to freeze a cohort for payback review (and when not to)

measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality

measurement
·1 min read