Incrementality testing for hybrid-casual games that works
By UA Ledger staff — Archive date: 5 min read

Incrementality testing for hybrid-casual games needs a different design than a standard holdout, because sessions and monetisation are both fast.
A holdout test built for a mid-core RPG does not transfer cleanly to a hybrid-casual title. Run it unmodified and you get one of the more common ways a hybrid-casual studio ends up distrusting incrementality testing altogether, concluding the method itself is unreliable when the actual problem was borrowing a design built for a different genre's engagement pattern. As I argued in "Incrementality Testing for Mobile Games You Can Afford," the spreadsheet-level version of a holdout is within reach for almost any team. What that piece didn't cover is how hybrid-casual's short sessions and blended IAP-plus-ad monetisation, together with its fast churn curves, change what the test needs to measure and how quickly you can read it. Get those adjustments wrong and the team walks away from a method that would have worked fine with the right design.
Why hybrid-casual complicates a standard holdout
A standard user-level or geo holdout assumes a reasonably stable gap between exposure and the outcome you measure (typically a purchase or a multi-day retention point). Hybrid-casual compresses both ends of that assumption. Sessions are short and frequent; monetisation events happen quickly for the users who convert at all; and a large share of revenue comes from ads rather than IAP, which most holdout designs built around purchase events never account for. A test that only tracks IAP conversion in a hybrid-casual holdout measures a minority of the game's actual revenue signal.
Session frequency and payback windows change the design
Because hybrid-casual users churn faster than mid-core users, the test window has to close faster too, or users who were never going to stick around, whichever arm they landed in, will dilute the read. A payback window built for a 90-day mid-core cohort is too slow for a genre where most of the meaningful retention and early monetisation signal shows up inside the first seven to fourteen days. So run the holdout for a shorter and tighter window and accept a smaller but faster-arriving sample. Stretching a hybrid-casual test to match a mid-core team's cadence is habit. Not design.
A worked example specific to hybrid-casual
Take a hypothetical hybrid-casual puzzle title running a network holdout across ten comparable geo clusters, with a blended revenue metric combining IAP and ad revenue per user rather than IAP alone. Over a two-week test window, the five clusters with the channel paused show blended revenue per user roughly 9% lower than the five clusters where spend continued, driven roughly two-thirds by lower ad revenue from reduced session volume and one-third by lower IAP conversion. A team that only tracked IAP here would have measured a smaller, less representative slice of the actual incremental effect. It would likely have undervalued the channel relative to its true contribution.
Reading results when monetisation is mixed
The practical checklist for a hybrid-casual incrementality test:
- Define the outcome metric as blended revenue per user (IAP plus ad revenue), not IAP alone, before the test starts.
- Shorten the test window to match the genre's actual retention curve, typically seven to fourteen days rather than a mid-core team's default of thirty or ninety.
- Segment the read by whether the lift shows up more in ad revenue (driven by session volume) or IAP (driven by purchase conversion), since the two point to different follow-up actions.
- Re-run the test after any major content update or live event, since hybrid-casual engagement patterns shift faster around content cadence than mid-core patterns do.
Adding ad revenue into the outcome metric solves the under-measurement problem. It also introduces noise. Ad revenue per user is itself sensitive to session volume and to mediation waterfall changes, and to fill-rate fluctuations that have nothing to do with the channel under test; a mid-core team's holdout design never has to deal with any of that. A geo cluster that happens to see a mediation partner's fill rate dip during the test window will show a revenue movement the holdout design will pin on the channel unless you check that confound separately. Pull mediation-level fill rate and eCPM data for the test window across both arms before finalising a read. Treat the result as suspect if fill-rate patterns diverged meaningfully between the held-out and active clusters for reasons unrelated to the channel itself.
Running the test alongside a live content calendar
Hybrid-casual titles rarely go two weeks without a live event or a content drop, or some seasonal push, and a holdout that ignores the content calendar risks measuring the event's effect rather than the channel's. The safer approach is to schedule the test window inside a stretch of the calendar with no major content changes planned, or, where that isn't possible, to run the event identically across all geo clusters in the test rather than staggering it by region. A team that lets a regional content rollout overlap with only half of its test clusters cannot separate the channel's incremental effect from the event's. The two effects are unlikely to be similar in size.
What this changes about channel decisions
A channel that looks marginal on IAP incrementality alone can look considerably stronger once ad revenue enters a hybrid-casual test, and the reverse holds: a channel driving IAP-heavy users who barely touch ad placements can look weaker than its attributed IAP numbers suggest once blended revenue is the yardstick. Neither result is available to a team still running the mid-core version of this test on a hybrid-casual title. Get the outcome metric and the test window right before the first read. That is the difference between a holdout that changes a budget decision and one that gets filed away as inconclusive.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·1 min read
Web-shop attributed revenue in MMP vs payment-provider settlements
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read