Incrementality Testing for Mobile Games You Can Afford
By Isaac Turner, Measurement Editor — Archive date: 6 min read
View author profile
Incrementality testing mobile games doesn't need a data science team. Holdout groups and geo splits can start on a spreadsheet and a modest budget.
Most UA teams that say they can't afford incrementality testing mobile games actually mean they can't afford the version vendors sell them: a dedicated measurement partner, a data science hire and a quarter of calendar time before the first read. That version exists and it works. It isn't the only version. A holdout built in a spreadsheet, run against a channel that already has enough volume to split, tells a studio more about what its paid spend is actually buying than another month of blended ROAS dashboards.
Why blended metrics keep lying
The problem incrementality testing solves is specific: attributed installs and revenue tell you what a channel claims credit for, not what it caused. A user who would have installed anyway, found through organic search or a friend's recommendation, still gets stamped as a paid conversion if an ad happened to be the last touch an MMP saw. Sensor Tower's State of Mobile 2025 report put total 2024 mobile gaming revenue at roughly $81 billion, up about 4% year on year, a market growing slowly enough that studios can't assume rising spend automatically buys rising incremental volume.
Every dollar spent on installs that would have arrived anyway is a dollar not spent on the installs that wouldn't have.
Attribution windows compound the problem on iOS, since SKAdNetwork models rather than observes much of the conversion path, and AdAttributionKit does the same. A channel with generous view-through windows and broad targeting will always look efficient in a blended report, because it claims credit liberally. Incrementality testing asks a narrower, harder question. If this channel's spend went to zero tomorrow, how many fewer installs would the game actually see the following week?
A design a small team can run
The standard approach, a public service announcement (PSA) or ghost bidding holdout, withholds a randomly selected slice of the addressable audience from ads on one channel while spend continues normally against the rest. The gap in organic-plus-other-channel conversion between the held-out group and the exposed group, adjusted for group size, is the incremental effect. Meta offers some form of built-in lift study tooling for this; so do Google and TikTok. None of it requires custom engineering.
The scaled-down version that fits a smaller budget is a geo holdout instead of a user-level PSA test. Pick a set of comparable regions or media markets, matched on historical install volume and revenue per user, and pause one channel's spend in half of them for two to four weeks while running normally in the rest. Then compare the change in blended install volume across both groups. It's a blunter instrument than a randomised user-level test, immune to some of the leakage that can affect audience-level splits, and it needs nothing more than a spend toggle plus a spreadsheet with weekly install counts by region.
A worked example
Take a hypothetical mid-size studio spending $40,000 a month on a UA channel across ten comparable geo clusters, and have it pause spend in five of them for three weeks while continuing normally in the other five. Before the test, all ten clusters tracked within 5% of each other on organic-plus-other installs per week. During the pause, the five held-out clusters show organic-plus-other installs roughly flat, while the five active clusters show a 12% increase over the same baseline. That 12% gap is the defensible estimate of what the channel is actually adding; the channel's attributed install count is not. If the channel's attributed installs for the active clusters implied a 40% contribution to total volume, and the test shows a 12% incremental lift, roughly two-thirds of the attributed credit was installs that would have happened anyway.
Reading incrementality testing results honestly
A geo holdout this size won't produce a publication-grade confidence interval, so don't treat it as one. As this column argued in "Web Shop Attribution Basics as Link-Outs Go Live in the EU," measurement built for a court case and measurement built for a Monday decision are different products, and demanding the first when only the second will do just means never shipping a test. The right bar for a geo holdout is whether the direction and rough magnitude of the result would change a spend decision. Peer review is not the standard.
Run the test again a quarter later before treating any single result as durable. Seasonality moves the true incremental share of a channel over time; so do competitive spend shifts and creative refreshes. A studio that runs one holdout in Q1 and reallocates budget on it for the rest of the year is substituting one untested assumption for another.
What to do with the number
Once a rough incremental share exists for a channel, the useful next step is comparing it against the channel's share of the blended budget. A channel that gets 30% of spend but tested at 12% incremental lift relative to its attributed volume is a candidate for a controlled spend reduction, tested the same way, before the next budget cycle. A channel that under-claims credit relative to its measured lift, which does happen, particularly with lower-funnel formats that convert users already close to installing, is a case for cautious reallocation upward.
None of this requires the incrementality-as-a-service packages some MMPs now sell as an add-on. Those tools automate the geo matching along with the statistical testing and the reporting, and they're worth paying for once a studio is running holdouts often enough that the manual version becomes the bottleneck. For a team running its first test this quarter, the spreadsheet version answers the only question that matters before the tooling question: has blended ROAS been overstating what the top channel actually buys? Answer that honestly first. The case for paying someone else to automate the process then becomes a budgeting decision rather than a leap of faith.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·1 min read
Web-shop attributed revenue in MMP vs payment-provider settlements
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read