Incrementality tests you can afford: geo holdouts, ghost ads, and what breaks

Analysis

By Isaac Turner, Measurement Editor — Archive date: 3 min read

Revised

View author profile
Abstract map regions in contrasting shades representing a geo holdout test

A practical guide to running geo holdout and ghost ad incrementality tests on a mid-size UA budget, including where each design breaks down.

The problem with attribution

Attribution assigns credit under a defined measurement rule. Incrementality asks what would have happened without the intervention. Those are different questions. Time-based regression models a counterfactual using historical treatment and control data. Experimental designs and causal models require explicit assumptions about the comparison group; a dashboard attribution total alone does not identify the counterfactual. This guide explains practical design questions, rather than promising that one method or fixed test duration can certify a campaign.

Two designs sit within reach of a mid-size team, and both break in ways worth knowing before you commit budget to either.

Geo holdouts

A geo holdout splits comparable markets or regions into a test group that keeps normal paid spend and a control group where you pause or cut it, then compares organic and total install or revenue trends between the two over a fixed window. Feasibility depends on whether spend can be changed in the selected markets and whether the available data supports a consistent outcome measure. Better still, the logic survives a conversation with a finance stakeholder who has never heard the word incrementality.

What breaks it is market comparability. Two regions rarely share the same seasonal curve or the same competitive intensity, their organic baselines drift apart for reasons nobody logs, and a holdout run over too short a window will read ordinary market noise as a treatment effect. Small markets break it too. Install volumes there sit too low for the difference between test and control to separate from noise at any confidence worth reporting. Choose duration and market allocation from baseline variance, expected effect, conversion lag and a power analysis. A longer run does not repair a systematically incomparable control.

Ghost ads and PSA tests

A ghost-ad experiment records ad opportunities where the advertiser would have served an impression to a control user, while excluding that advertiser from the actual control delivery. Another eligible ad can be shown. The original Ghost Ads research, introduction and methodology distinguishes this exposure logging from public-service-announcement controls, which substitute a different creative. A blank placeholder is not the defining mechanism.

Platform support and experimental implementation determine whether a ghost-ad approach is available. Ask what the platform logs, how treatment assignment is defined and what the estimator measures. Power depends on the design and outcome variance; do not assume a universal minimum spend. A result applies to the tested population, period and treatment and does not automatically transfer to another channel.

What to actually run

Start by writing the estimand: incremental installs, revenue or contribution over a stated horizon. Check historical variance and geographic spillovers, define the minimum effect worth detecting and agree the analysis before changing spend. Pilot the data pipeline first. A geo holdout and a platform experiment answer different scoped questions; do not treat either as universal truth or prescribe a six-week run without the supporting assumptions.

Related reading

Corrections & updates

  • Correction (): Corrected ghost ads versus PSA controls; removed categorical modelling and universal-duration prescriptions.

Featured

Related posts

measurement

·

2 min read

When to freeze a cohort for payback review (and when not to)

measurement

media buying

·

1 min read

Using predicted LTV in bids: disclosure checklist for the UA team

measurement

media buying

·

1 min read

Blended ROAS targets that hide channel failure in F2P portfolios

measurement

·

1 min read

Web-shop purchaser quality vs store IAP purchaser quality

More from the Measurement desk

measurement

·

2 min read

Airbridge adds Amazon Ads as an app measurement channel

measurement

·

1 min read

Web-shop attributed revenue in MMP vs payment-provider settlements

measurement

media buying

·

2 min read

Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)

measurement

media buying

·

1 min read

View-through attribution windows on F2P rewarded and interstitial traffic