Scale a holdout outcome rate before subtracting conversions
Analysis
By Isaac Turner, Measurement Editor — 2 min read
View author profile
A treatment group twice the size of its holdout should not be compared by raw conversion counts. A reproducible synthetic example with editable inputs and explicit limits.
A treatment group twice the size of its holdout should not be compared by raw conversion counts. This lab translates the holdout rate into an expected count at the treatment population size, then calculates a descriptive excess. Causal interpretation still depends on the experiment.
Define the calculation before using it
Use eligible assigned people in each group and a shared conversion definition. Divide holdout conversions by holdout population, multiply that rate by treatment population, then subtract the resulting expected count from treatment conversions. Keep the rate difference visible alongside the count difference. Do not use reached users in one arm and assigned users in the other.
Work through the synthetic example
The example has 1,000 treated assignments with 80 conversions and 500 holdout assignments with 30. The holdout rate implies 60 conversions at the treatment size, leaving an excess of 20 and a two-percentage-point rate difference. Those are synthetic descriptive results, not a measured campaign lift or a confidence interval.
| Output | Worked-example result |
|---|---|
| Scaled holdout conversions | 60.0000 |
| Descriptive excess conversions | 20.0000 |
| Rate difference percentage points | 2.0000 |
Use the artifact and preserve its assumptions
Open the editable calculator to change the inputs and inspect the sensitivity view. The CSV records synthetic inputs and expected outputs; the JSON fixture keeps the equations available for reproduction. These calculations have been checked against the stated example. No measured campaign data is included.
The sensitivity rows vary only cc by 20% below and above the entered value. They are scenarios, not confidence limits or a forecast distribution. A row outside the model’s constraints is labelled rather than turned into a plausible-looking result. Save the chosen inputs with the decision so another reader can distinguish a changed assumption from a changed formula.
Evidence and limits
NIST discusses confidence intervals for a binomial proportion and cautions about approximation with small samples. Our worksheet states its assumptions; a descriptive rate or count contrast alone does not establish statistical certainty. See Confidence intervals for a proportion, especially “Confidence intervals; small-sample exact intervals”.
Random assignment, comparable observation and limited spillover are unverified. The worksheet does not model sampling uncertainty, noncompliance, household clustering or interactions with concurrent campaigns.
Background: Geo-lift tests: a practical guide for UA teams and Incrementality tests you can afford: geo holdouts, ghost ads, and what breaks. These existing articles provide context; the present calculation does not verify every archived claim.
Featured
Related posts
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read
View-through attribution windows on F2P rewarded and interstitial traffic
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
measurement
·1 min read