Check an experiment’s assignment counts before reading its uplift
Analysis
By Isaac Turner, Measurement Editor — 2 min read
View author profile
An experiment can produce an attractive uplift while its assignment counts contradict the intended split. A reproducible synthetic example with editable inputs and explicit limits.
An experiment can produce an attractive uplift while its assignment counts contradict the intended split. This lab calculates a two-group chi-square diagnostic for the allocation counts. It is an initial integrity check, not a declaration that an experiment is valid or invalid on its own.
Define the calculation before using it
Use the randomised assignment population rather than only people who completed a downstream event. Derive each expected count from the total assignments and planned split, then sum squared observed-minus-expected differences divided by expected counts. The worksheet reports the statistic without a p-value or automatic pass threshold. Predefine the statistical review and investigate the actual assignment mechanism.
Work through the synthetic example
A planned 50:50 allocation with 550 assignments in A and 450 in B expects 500 per group. The two contributions are five each, giving a statistic of ten. The number is reproducible; its operational meaning depends on the design, observation completeness and whether the counts represent independent assignments.
| Output | Worked-example result |
|---|---|
| Expected A | 500.0000 |
| Chi-square statistic | 10.0000 |
Use the artifact and preserve its assumptions
Open the editable calculator to change the inputs and inspect the sensitivity view. The CSV records synthetic inputs and expected outputs; the JSON fixture keeps the equations available for reproduction. These calculations have been checked against the stated example. No measured campaign data is included.
The sensitivity rows vary only a by 20% below and above the entered value. They are scenarios, not confidence limits or a forecast distribution. A row outside the model’s constraints is labelled rather than turned into a plausible-looking result. Save the chosen inputs with the decision so another reader can distinguish a changed assumption from a changed formula.
Evidence and limits
NIST describes the chi-square statistic using observed and expected counts. This lab applies that arithmetic to a two-group assignment diagnostic without providing an automatic significance verdict. See Chi-Square Goodness-of-Fit Test, especially “Test Statistic; Assumptions”.
Very small expected cells, clustered assignment and repeated checks need specialist treatment. A balanced allocation also cannot establish correct event logging, absence of spillover or a causal effect on revenue.
Background: UA metrics explained: CPI, ROAS, LTV and payback and Reading an MMP dashboard without fooling yourself. These existing articles provide context; the present calculation does not verify every archived claim.
Featured
Related posts

measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team

measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios

measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)

measurement
media buying
·1 min read
View-through attribution windows on F2P rewarded and interstitial traffic
More from the Measurement desk

measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel

measurement
·2 min read
When to freeze a cohort for payback review (and when not to)

measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality

measurement
·1 min read