Soft launch UA plans and kill criteria
By Isaac Turner, Measurement Editor — Archive date: 5 min read
View author profile
A working framework for soft launch UA: country choice, spend ramps, retention gates and the kill criteria that stop sunk-cost extensions.
Most soft launch plans have a country list and a budget. Very few have a written rule for when the test ends. That gap is where studios lose months, because a soft launch without a pre-agreed exit condition tends to become a slow argument rather than a decision, and the team that built the game is rarely the team best placed to call time on it.
Country selection is a measurement choice, not a cost choice
The instinct is to pick cheap CPI markets: Philippines, Vietnam, parts of Latin America. That keeps the burn low, but it also means the retention and monetisation curves you collect describe a different player than your target market. A studio aiming at US or Western European IAP revenue needs at least one soft launch country that shares payer behaviour with the target, even if CPI there is three to five times higher than in a pure volume market. Australia, Canada and the Nordics remain the standard proxies for US-like spend behaviour at a fraction of full US media cost. Pair one proxy market with one or two cheap volume markets for statistical power on early funnel metrics, and treat the two data sets as answering different questions rather than averaging them together.
The spend ramp should be staged, not front-loaded
A common failure mode is spending the full soft launch budget in the first two weeks to get a fast read. That produces a fast read of the wrong thing, because early cohorts are disproportionately made up of players reached through the cheapest, most efficient inventory, not a representative sample of what full launch spend will buy. A staged ramp, roughly a quarter of budget in each of four phases, each phase widening the channel and creative mix, gives a truer picture of what marginal spend looks like as it moves down the demand curve. It also gives the team room to swap out weak creative before the whole budget is spent chasing it.
Metrics that should actually decide the outcome
Day one and day seven retention get the attention, but they are weak predictors on their own. Day thirty retention combined with early payer conversion and the first realistic look at day thirty or day sixty return on ad spend is a better basis for a call, and that means a soft launch that ends before day thirty has usually not run long enough to answer the question it was set up to ask. The specific gates worth setting before the test starts:
- A minimum day one retention threshold below which no amount of later optimisation has historically saved a title in the genre.
- A day thirty ROAS band, expressed against payback target, not against an arbitrary benchmark from a different genre.
- A minimum sample size per cohort before any retention or monetisation number is treated as a signal rather than noise.
Genre matters enormously here. A puzzle title and a mid-core strategy game have different acceptable retention curves and different payback horizons, and importing a threshold from the wrong genre comparison is one of the more common soft launch mistakes.
Building the kill criteria before you need them
The reason kill criteria fail in practice is not that teams disagree on the numbers when the moment comes. It is that nobody wrote the numbers down before the test started, so the moment becomes a negotiation between a UA lead who wants to stop and a studio that has spent a year on the game. Writing the thresholds into the soft launch plan itself, signed off by whoever controls the budget, before a single dollar is spent, removes most of that friction. It also protects the UA team from being blamed for a call that was actually made collectively months earlier.
A kill criterion is not just a floor. It should also specify what an extension requires: a named metric that improved since the last review, and by how much, not a general sense that "the team is close." Extensions granted on vibes are how six-week soft launches become six-month ones with no cleaner data at the end than at the start.
What to do with a borderline result
Most soft launches do not fail cleanly. They land in the band between clear kill and clear scale, and that band is where the sunk-cost problem is worst. The useful move here is to separate the retention question from the monetisation question. A game with strong retention and weak monetisation is a design and pricing problem that a UA team cannot fix with better targeting, and continuing to spend against it will not change the underlying curve. A game with adequate retention and adequate but not exciting monetisation is often a genuine judgment call, and that is where a second, narrower test, one variable changed, one market, a fixed budget and a fixed end date, earns its cost far better than an open-ended extension of the original test.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
market intelligence
measurement
·1 min read
Retention percentile charts for F2P: what GameAnalytics-style cuts miss
market intelligence
measurement
·2 min read
Sensor Tower’s $82 billion figure is an IAP measure
market intelligence
measurement
·2 min read
The 52 billion download headline crosses platforms
market intelligence
measurement
·2 min read
One casual-gaming report contains several data populations
More from the Market Intelligence desk
market intelligence
media buying
·1 min read
Lunar New Year UA gates from an official calendar, not a CPI myth
market intelligence
media buying
·2 min read
WeChat Mini Games upgrades 2026 IAP / virtual-payment incentives for debut titles
market intelligence
media buying
·2 min read
WeChat Mini Games IAA incentives from 20 August 2026: 3-minute lifetime, 180-day option
creative strategy
market intelligence
·1 min read