Holiday spend incrementality holdout tests done cheaply
By Isaac Turner, Measurement Editor — Archive date: 6 min read
View author profile
A holiday spend incrementality holdout test doesn't need a big budget, and running one during the Q4 ramp tells you more than any blended ROAS number.
Every UA team is about to spend more than it has spent all year, and almost none of them will know, once the quarter closes, how much of that extra spend actually bought incremental installs. A holiday spend incrementality holdout test answers that question directly, and it costs less to run than most teams assume. The excuse for skipping one, that Q4 is too busy to add a testing workstream on top of an already stretched media plan, gets the priority backwards. Q4 is exactly the period when the answer is worth the most, because it is the period when the largest share of the annual budget is being deployed against the least reliable read of what that spend is buying.
Why blended ROAS lies hardest in Q4
Blended return on ad spend gets noisier as spend increases, not steadier, because a rising tide of organic and brand-driven installs during a high-attention shopping period gets attributed to paid campaigns that happen to be running at the same time. A campaign that looks like it is performing brilliantly in the second week of December may simply be riding a seasonal lift in organic discovery that would have happened with or without the spend increase. Without a holdout, there is no way to separate the two, and the gap between what a dashboard shows and what spend actually caused tends to be largest precisely when the stakes are highest.
The geo holdout, sized for Q4
A geo-based holdout test is the cheapest reliable design available, and it does not require pausing spend everywhere to get a clean read.
- Pick a small set of comparable markets, ideally five to ten, matched on size and recent performance rather than picked for convenience.
- Hold spend flat or paused in a subset representing 10% to 20% of those markets' typical volume, and increase spend normally everywhere else in the test group.
- Run the test for the full length of the planned spend increase, not a truncated week, since holiday installs and their downstream monetisation take time to show up.
- Compare organic and total install trends between the holdout and matched markets, not just paid-attributed installs, since the whole point is catching what a dashboard would otherwise credit to paid.
A worked example, with hypothetical numbers
Take a hypothetical hybrid-casual publisher planning to lift UA spend from $200,000 to $350,000 a week across its top ten markets for the four weeks around Black Friday and Cyber Monday. Holding two matched markets flat at their normal $15,000-a-week spend level while the other eight scale up gives a holdout covering roughly 15% of typical volume, small enough not to meaningfully dent the quarter's install total but large enough to produce a readable signal. If the eight scaled markets show a 60% lift in total installs against the two holdout markets showing only a 10% lift from organic seasonality alone, that gap, roughly 50 percentage points, is a defensible estimate of what the spend increase actually bought. If the gap were closer to 15 or 20 points, that would say the spend increase is mostly riding a seasonal wave the publisher would have caught anyway. These numbers are illustrative only, built to show the shape of the comparison, not a benchmark to import into your own model.
The pitfalls that ruin a holiday holdout
A holdout test run casually during Q4 is easy to contaminate, and the two most common mistakes both stem from the same source: the pressure to keep every dollar working during the highest-spend weeks of the year.
- Letting a "quiet" holdout market receive spillover spend: a campaign manager under pressure to hit a volume target will sometimes quietly extend budget into a holdout market mid-test, which invalidates the comparison without anyone flagging it as a decision. Lock the holdout list with whoever controls budget allocation before the test starts, not as a verbal agreement that can slip under deadline pressure.
- Confusing a short lull with a finished test: holiday installs and their downstream revenue take longer to mature than a normal week's cohort, since gifted downloads and delayed first sessions are common in gift-giving markets. Reading a holdout result after only a few days risks catching a partial signal rather than the full incremental lift.
- Sizing the holdout too small to be readable: a holdout covering only 2% or 3% of volume may not produce a statistically distinguishable gap from noise, particularly in markets with naturally volatile organic install rates. The 10% to 20% range suggested above is a starting point, not a fixed rule, and should be sized larger in noisier markets.
What this changes for the January review
The value of a holiday spend incrementality holdout test is not just the number it produces in December. It sets a baseline for the January post-mortem that every finance stakeholder will ask for anyway, and it means that conversation starts from an incremental figure instead of a blended one that overstates the quarter's paid performance. As we discussed in Incrementality Testing for Mobile Games You Can Afford, the barrier to running these tests has always been perceived cost and complexity more than actual cost, and a single well-sized geo holdout during the one period of the year with the most spend at stake is the cheapest version of that argument to make to a sceptical CFO. Build the holdout into the media plan before the spend increase starts, not as a retrofit once December numbers already look confusing.
Holiday spend incrementality holdout test: what to watch next
Two things are worth checking as the test runs, not just at the end. Watch whether the holdout markets show any unexpected demand shock unrelated to UA, such as a competitor's own holiday campaign concentrated in one of the chosen regions, since that would contaminate the comparison independent of anything the publisher's own team did wrong. And watch whether finance is already planning to use the December blended ROAS number in a year-end report before the holdout's incremental figure is ready, since a mismatch between the two numbers appearing in different documents at different times is a much easier conversation to have proactively than after someone has already presented the wrong one to leadership.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read
View-through attribution windows on F2P rewarded and interstitial traffic
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
measurement
·1 min read