Geo-lift tests: a practical guide for UA teams
By Isaac Turner, Measurement Editor — Archive date: 5 min read
View author profile
How to design a geo-lift holdout for mobile UA: sizing for power, the minimum spend needed, and the three ways teams misread the result.
A geo-lift test answers a question no attribution dashboard can: what happens to installs and revenue in a region when you simply stop advertising there. Withhold spend in a set of regions, keep spending as normal in the rest, and compare what actually happened in the holdout against what your model expected to happen with no ads at all. The difference is your incremental lift, uncontaminated by the click-attribution logic that credits an install to whichever touchpoint happened to sit closest to it.
Choosing the geo unit
The unit needs to be small enough that you have many of them, so the test has statistical power, but large enough that media buying can plausibly be turned off there without leaking spend across the boundary. In the US, designated market areas work well because Nielsen-style DMA boundaries roughly track how local media actually reaches people, and most ad platforms let you exclude a DMA list directly. Outside the US, country-level holdouts are the practical default for a mobile game studio, since sub-national ad exclusion is patchier to execute cleanly across networks, though city-tier holdouts work in markets like India or Brazil where local buying teams already segment by tier.
Whatever the unit, resist the temptation to hand-pick which regions go into the holdout based on which ones look convenient. Match holdout and test regions on pre-period installs, revenue per install and organic share before you randomise, then randomise within matched pairs. A holdout that happens to contain your three weakest markets will always show a smaller lift than reality, because those markets were never going to respond much to any spend.
Sizing for power
Before running anything, estimate how large a lift you could actually detect given your regions' baseline variance. A rough approach: pull twelve months of weekly installs per candidate region, compute the coefficient of variation, and run a power calculation assuming you want to detect the lift size your finance team actually cares about, typically somewhere between 10 and 30 percent depending on the channel's spend share. Teams that skip this step commonly run a test with far too few regions or too short a duration to detect anything short of a very large effect, then wrongly conclude the channel is not incremental when the real finding is that the test had no power to find out either way.
As a working rule, four to six weeks is close to the minimum test duration for a mobile game with any meaningful sales cycle or event calendar, since a shorter window gets swamped by day-to-day noise and by any in-game event that happens to land inside it. The minimum spend threshold to bother testing at all is whatever level makes the expected lift larger than your regions' typical week-to-week swing, which for a mid-sized studio channel is usually somewhere in the low tens of thousands of pounds a week, not a fixed industry number.
Reading the result
The comparison is not holdout versus test region in absolute terms, it is holdout versus a counterfactual built from the holdout's own pre-period trend plus whatever the matched test regions did during the test window. Most geo-lift tooling, whether a vendor platform or an in-house synthetic control script, builds this counterfactual automatically, but a UA lead should be able to eyeball the underlying chart and sanity-check that the pre-period lines track each other closely before trusting the post-period gap.
Three ways teams fool themselves
The first mistake is treating a single test as durable. A channel's incrementality shifts with your organic strength, your competitive set and the season, so a test run once during a slow month gets quietly cited as gospel a year later. Re-test on a cadence, not once.
The second is contamination across the holdout boundary. Ad platforms with broad geo-targeting, especially anything running lookalike or interest-based expansion, can leak impressions into a nominally excluded region. Check delivery reports for the holdout regions during the test, not just after it, and be prepared to tighten exclusions mid-flight.
The third is confusing a null result for a bad channel. A geo-lift test that shows no detectable lift in a channel that is genuinely small relative to your regions' baseline noise is not evidence the channel does nothing, it is evidence the test was underpowered. Pair every result with the confidence interval, not just the point estimate, before it goes into a budget decision, and treat a wide interval around zero as "inconclusive," not "cut this channel."
Geo-lift is one of the few incrementality methods that gives a UA team a defensible answer finance will accept, because it does not depend on the same attribution logic the finance team is already sceptical of. That is also why it is worth running properly rather than as a box-ticking exercise once a year.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·1 min read
Web-shop attributed revenue in MMP vs payment-provider settlements
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read