Scale a holdout outcome rate before subtracting conversions

Analysis

By Isaac Turner, Measurement Editor2 min read

View author profile
Scale a holdout outcome rate before subtracting conversions

A treatment group twice the size of its holdout should not be compared by raw conversion counts. A reproducible synthetic example with editable inputs and explicit limits.

A treatment group twice the size of its holdout should not be compared by raw conversion counts. This lab translates the holdout rate into an expected count at the treatment population size, then calculates a descriptive excess. Causal interpretation still depends on the experiment.

Define the calculation before using it

Use eligible assigned people in each group and a shared conversion definition. Divide holdout conversions by holdout population, multiply that rate by treatment population, then subtract the resulting expected count from treatment conversions. Keep the rate difference visible alongside the count difference. Do not use reached users in one arm and assigned users in the other.

Work through the synthetic example

The example has 1,000 treated assignments with 80 conversions and 500 holdout assignments with 30. The holdout rate implies 60 conversions at the treatment size, leaving an excess of 20 and a two-percentage-point rate difference. Those are synthetic descriptive results, not a measured campaign lift or a confidence interval.

Reference table
OutputWorked-example result
Scaled holdout conversions60.0000
Descriptive excess conversions20.0000
Rate difference percentage points2.0000

Use the artifact and preserve its assumptions

Open the editable calculator to change the inputs and inspect the sensitivity view. The CSV records synthetic inputs and expected outputs; the JSON fixture keeps the equations available for reproduction. These calculations have been checked against the stated example. No measured campaign data is included.

The sensitivity rows vary only cc by 20% below and above the entered value. They are scenarios, not confidence limits or a forecast distribution. A row outside the model’s constraints is labelled rather than turned into a plausible-looking result. Save the chosen inputs with the decision so another reader can distinguish a changed assumption from a changed formula.

Evidence and limits

NIST discusses confidence intervals for a binomial proportion and cautions about approximation with small samples. Our worksheet states its assumptions; a descriptive rate or count contrast alone does not establish statistical certainty. See Confidence intervals for a proportion, especially “Confidence intervals; small-sample exact intervals”.

Random assignment, comparable observation and limited spillover are unverified. The worksheet does not model sampling uncertainty, noncompliance, household clustering or interactions with concurrent campaigns.

Background: Geo-lift tests: a practical guide for UA teams and Incrementality tests you can afford: geo holdouts, ghost ads, and what breaks. These existing articles provide context; the present calculation does not verify every archived claim.

Featured

Related posts

measurement

media buying

·

1 min read

Using predicted LTV in bids: disclosure checklist for the UA team

measurement

media buying

·

1 min read

Blended ROAS targets that hide channel failure in F2P portfolios

measurement

media buying

·

2 min read

Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)

measurement

media buying

·

1 min read

View-through attribution windows on F2P rewarded and interstitial traffic

More from the Measurement desk

measurement

·

2 min read

Airbridge adds Amazon Ads as an app measurement channel

measurement

·

2 min read

When to freeze a cohort for payback review (and when not to)

measurement

·

1 min read

Web-shop purchaser quality vs store IAP purchaser quality

measurement

·

1 min read

Web-shop attributed revenue in MMP vs payment-provider settlements