Incrementality Testing for Mobile Games You Can Afford

By Isaac Turner, Measurement Editor — Archive date: 6 min read

View author profile
Abstract split panel showing one lit and one dark cohort of user icons

Incrementality testing mobile games doesn't need a data science team. Holdout groups and geo splits can start on a spreadsheet and a modest budget.

Most UA teams that say they can't afford incrementality testing mobile games actually mean they can't afford the version vendors sell them: a dedicated measurement partner, a data science hire and a quarter of calendar time before the first read. That version exists and it works. It isn't the only version. A holdout built in a spreadsheet, run against a channel that already has enough volume to split, tells a studio more about what its paid spend is actually buying than another month of blended ROAS dashboards.

Why blended metrics keep lying

The problem incrementality testing solves is specific: attributed installs and revenue tell you what a channel claims credit for, not what it caused. A user who would have installed anyway, found through organic search or a friend's recommendation, still gets stamped as a paid conversion if an ad happened to be the last touch an MMP saw. Sensor Tower's State of Mobile 2025 report put total 2024 mobile gaming revenue at roughly $81 billion, up about 4% year on year, a market growing slowly enough that studios can't assume rising spend automatically buys rising incremental volume.

Every dollar spent on installs that would have arrived anyway is a dollar not spent on the installs that wouldn't have.

Attribution windows compound the problem on iOS, since SKAdNetwork models rather than observes much of the conversion path, and AdAttributionKit does the same. A channel with generous view-through windows and broad targeting will always look efficient in a blended report, because it claims credit liberally. Incrementality testing asks a narrower, harder question. If this channel's spend went to zero tomorrow, how many fewer installs would the game actually see the following week?

A design a small team can run

The standard approach, a public service announcement (PSA) or ghost bidding holdout, withholds a randomly selected slice of the addressable audience from ads on one channel while spend continues normally against the rest. The gap in organic-plus-other-channel conversion between the held-out group and the exposed group, adjusted for group size, is the incremental effect. Meta offers some form of built-in lift study tooling for this; so do Google and TikTok. None of it requires custom engineering.

The scaled-down version that fits a smaller budget is a geo holdout instead of a user-level PSA test. Pick a set of comparable regions or media markets, matched on historical install volume and revenue per user, and pause one channel's spend in half of them for two to four weeks while running normally in the rest. Then compare the change in blended install volume across both groups. It's a blunter instrument than a randomised user-level test, immune to some of the leakage that can affect audience-level splits, and it needs nothing more than a spend toggle plus a spreadsheet with weekly install counts by region.

A worked example

Take a hypothetical mid-size studio spending $40,000 a month on a UA channel across ten comparable geo clusters, and have it pause spend in five of them for three weeks while continuing normally in the other five. Before the test, all ten clusters tracked within 5% of each other on organic-plus-other installs per week. During the pause, the five held-out clusters show organic-plus-other installs roughly flat, while the five active clusters show a 12% increase over the same baseline. That 12% gap is the defensible estimate of what the channel is actually adding; the channel's attributed install count is not. If the channel's attributed installs for the active clusters implied a 40% contribution to total volume, and the test shows a 12% incremental lift, roughly two-thirds of the attributed credit was installs that would have happened anyway.

Reading incrementality testing results honestly

A geo holdout this size won't produce a publication-grade confidence interval, so don't treat it as one. As this column argued in "Web Shop Attribution Basics as Link-Outs Go Live in the EU," measurement built for a court case and measurement built for a Monday decision are different products, and demanding the first when only the second will do just means never shipping a test. The right bar for a geo holdout is whether the direction and rough magnitude of the result would change a spend decision. Peer review is not the standard.

Run the test again a quarter later before treating any single result as durable. Seasonality moves the true incremental share of a channel over time; so do competitive spend shifts and creative refreshes. A studio that runs one holdout in Q1 and reallocates budget on it for the rest of the year is substituting one untested assumption for another.

What to do with the number

Once a rough incremental share exists for a channel, the useful next step is comparing it against the channel's share of the blended budget. A channel that gets 30% of spend but tested at 12% incremental lift relative to its attributed volume is a candidate for a controlled spend reduction, tested the same way, before the next budget cycle. A channel that under-claims credit relative to its measured lift, which does happen, particularly with lower-funnel formats that convert users already close to installing, is a case for cautious reallocation upward.

None of this requires the incrementality-as-a-service packages some MMPs now sell as an add-on. Those tools automate the geo matching along with the statistical testing and the reporting, and they're worth paying for once a studio is running holdouts often enough that the manual version becomes the bottleneck. For a team running its first test this quarter, the spreadsheet version answers the only question that matters before the tooling question: has blended ROAS been overstating what the top channel actually buys? Answer that honestly first. The case for paying someone else to automate the process then becomes a budgeting decision rather than a leap of faith.

Related archive reading

These articles provide related context and remain subject to their stated review status.

Featured

Related posts

measurement

·

2 min read

When to freeze a cohort for payback review (and when not to)

measurement

media buying

·

1 min read

Using predicted LTV in bids: disclosure checklist for the UA team

measurement

media buying

·

1 min read

Blended ROAS targets that hide channel failure in F2P portfolios

measurement

·

1 min read

Web-shop purchaser quality vs store IAP purchaser quality

More from the Measurement desk

measurement

·

2 min read

Airbridge adds Amazon Ads as an app measurement channel

measurement

·

1 min read

Web-shop attributed revenue in MMP vs payment-provider settlements

measurement

media buying

·

2 min read

Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)

measurement

media buying

·

1 min read

View-through attribution windows on F2P rewarded and interstitial traffic