Anatomy of a soft launch that should have been killed
By UA Ledger staff — Archive date: 6 min read

Soft launches survive too long because kill criteria are set on metrics the team can move with effort, not the metrics that reveal the market.
Soft launches rarely end with a decision. They end with a shipping date. The game the studio should have stopped in month three is still in test in month nine after hitting and re-hitting a sequence of targets, and by the time the honest conversation happens the studio has spent the money that a kill was meant to save.
The reason is not that the data is ambiguous. It is that the kill criteria almost always sit on metrics the team can move through effort, and a motivated team will move them. The tutorial tunes D1 retention; rewards tune session length. Neither tells you whether the market wants the game. The metrics that do, which are the cost to acquire a player at a volume that matters, the payer conversion of unaided cohorts, and the ratio of organic to paid installs, are the ones most soft launch plans treat as later-stage concerns. By the time anyone measures them, the exit has become politically impossible.
The two kinds of metric
Sort every number in a soft launch dashboard into two bins.
Movable metrics respond to work by the team. D1 and D3 retention, tutorial completion, early session count, first-purchase timing. They are legitimate targets for a team improving a game, and they are terrible kill criteria, because a kill criterion should test a hypothesis about the world, not the team's diligence.
Market metrics respond to what the audience is willing to do. CPI at a meaningful daily budget on the channels you will actually use at launch. Payer conversion and early ARPPU in cohorts that did not arrive on incentivised or heavily filtered traffic. The organic share of installs when paid is running. D30 and D60 retention shape, which the tutorial cannot reach.
A soft launch that ran on past its kill point almost always has a story that goes: the team flagged the market metrics as not yet reliable, the movable metrics improved sprint after sprint, and each improvement bought another two months. The team was not deceiving anyone; it was doing what the criteria rewarded.
Why the test geo lies
The choice of test market compounds the problem. The standard candidate list exists because installs there are cheap, and cheap installs let a small budget produce statistically comfortable cohort sizes. But cheapness distorts the very signal you're after. A CPI that looks acceptable at a few hundred installs a day in a low-cost market says nothing about the price of the marginal install in the United States at ten times that volume, where the launch business case will stand or fall.
Worse, at small budgets the network's optimiser is skimming. It finds the cheapest few percent of the available audience for your creative and serves those, so the CPI you observe is the price of the easiest users to convert, not the price of the users you need. This is the same mechanism that makes small network tests unreliable, and it is why a soft launch CPI deserves treatment as a floor rather than an estimate until you have pushed spend far enough that the CPI curve bends.
The incentive picture matters too. A studio's soft launch team gets paid to ship. A publisher's producer often carries a slate target. The publisher's portfolio logic, which is the correct logic for a portfolio, is that most titles fail and the winners pay for them, but portfolio logic only works if the failures fail quickly. The consolidation seen in hybrid-casual publishing this year, including Tripledot's completed purchase of Supersonic from Unity earlier this month, puts more titles under owners whose economics depend on fast kills. That is a reason to expect kill discipline to tighten among the large publishers and a reason for independent studios to worry that their own discipline is comparatively lax.
The criteria that would have worked
Write kill criteria as dated tests of market hypotheses. Each should be something the team cannot improve except by making the game more wanted.
- By a fixed date, CPI on the launch channel in one launch-representative geo at a daily budget large enough to bend the curve. The threshold comes from the launch business case, not from the test-market number.
- By a fixed date, payer conversion and early ARPPU in cohorts from that geo, excluding any incentivised or rewarded install traffic.
- By a fixed date, D30 retention measured, not projected from D7.
- At every review, the answer to one question: knowing what we know today, would we fund this game from scratch at this budget. If the answer is no, the fact that the money is already spent is not a reason to spend more.
The dates are the point. A criterion without a date is a hope. Slipping a date should be a formal decision recorded with the reason, and two slippages on the same criterion is itself a kill signal, because it means the team is dodging the market metric.
An illustrative case
Consider an invented mid-core game in month seven of soft launch, with numbers for arithmetic only. D1 retention has moved from 32 to 41 percent across five tutorial revisions. D7 has moved from 9 to 12. The team is proud, correctly. But D30 has held at around 3 percent throughout, payer conversion in the test geo is under 1 percent, and the one week the team pushed the budget to launch-representative levels in a tier-one market, CPI came in at roughly three times the figure in the business case before the team pulled back and called the test inconclusive.
Every movable metric improved. Every market metric said no, and the team measured the market metrics last and least, with the smallest budgets. Dated market criteria would have stopped this game in month four, saving roughly three months of a full team plus test spend. Under movable criteria it will ship, launch to a CPI it cannot afford, and quietly sunset within a year.
The trade-off nobody likes
Market criteria are expensive to measure. Bending the CPI curve in a tier-one geo costs real money on a game the result may kill. The temptation is to defer that spend until the game is better, which is exactly how a soft launch outlives its kill point. The honest position is that a fraction of the soft launch budget, call it a fifth, is the price of the kill decision itself, and it should go out early enough that the decision remains open.
Invite-only launches, discussed in "Supercell's mo.co and a Soft Launch Strategy for Games", solve a different problem: they protect a brand while the game remains unfinished. They do not remove the need to buy a market signal at scale before the launch plan locks.
The last thing a kill criterion needs is an owner who does not report into the team under test. A producer cannot kill their own project on schedule any more than a UA manager can call their own channel test a failure. Name the person who holds the date before the soft launch begins.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
market intelligence
measurement
·1 min read
Retention percentile charts for F2P: what GameAnalytics-style cuts miss
market intelligence
measurement
·2 min read
Sensor Tower’s $82 billion figure is an IAP measure
market intelligence
measurement
·2 min read
The 52 billion download headline crosses platforms
market intelligence
measurement
·2 min read
One casual-gaming report contains several data populations
More from the Market Intelligence desk
market intelligence
media buying
·1 min read
Lunar New Year UA gates from an official calendar, not a CPI myth
market intelligence
media buying
·2 min read
WeChat Mini Games upgrades 2026 IAP / virtual-payment incentives for debut titles
market intelligence
media buying
·2 min read
WeChat Mini Games IAA incentives from 20 August 2026: 3-minute lifetime, 180-day option
creative strategy
market intelligence
·1 min read