Anatomy of a store page test that moved nothing
By UA Ledger staff — Archive date: 6 min read

A flat store page test is rarely evidence that the store page does not matter. It is usually evidence that the test could not see the traffic that does.
A store page test that comes back flat is treated by most teams as a null result about the store page. It is more often a null result about the test. The store page does not have one conversion rate. It has as many conversion rates as there are traffic sources landing on it, and a test that averages across all of them is measuring the mix of visitors far more than the quality of the page.
That claim cuts against how most studios run store experiments, so the reasoning matters. A browsing visitor who arrives from a category chart, a visitor who arrives from a branded search, and a visitor who arrives from a rewarded playable have already made three different decisions before the page loads. The browser is undecided and reads. The searcher has decided and taps. The playable visitor has just spent thirty seconds performing the game and is checking that the page matches what they played. Screenshot variant B can help the first, be invisible to the second, and actively confuse the third. Averaged together, the result is a flat line and a wrong conclusion.
Why the average hides the effect
The mechanism is a composition problem. Suppose, as an illustrative example, that a page receives sixty per cent of its visits from paid campaigns, thirty from search, and ten from browse. Search visitors convert at a high and nearly fixed rate regardless of what the screenshots show, because the search itself was the decision. Paid visitors convert at a rate set largely by how well the page continues the creative they tapped. Browse visitors are the only group for whom the page is doing classic persuasion.
A test variant that improves browse conversion by a meaningful margin moves the blended figure by a fraction of that, because browse is a tenth of the traffic. A variant that improves paid conversion by matching a specific campaign's hook would move the blended figure more, but it would also degrade conversion for every other campaign's visitors, and those effects partly cancel. Either way the headline reads flat, and either way the team concludes that the store page is not a lever.
Both conclusions are wrong for the same reason. The page is a lever. The test lacked the resolution to see it pulled.
The second-order cost of a flat result
The direct cost of a badly designed store test is wasted experiment time. The indirect cost is larger and slower. A flat result tends to close the topic for a quarter or two. The team stops investing in store creative, the screenshots drift out of alignment with whatever the current winning ad concept is, and paid conversion decays in a way that is invisible because it is happening on the platform's side of the attribution boundary.
Meanwhile the media buyer sees cost per install creeping up and attributes it to auction pressure or creative fatigue. The correct diagnosis, that the store page no longer matches the ad, does not surface because the last store test said the store page did not matter. UA Ledger made the creative-side argument in The store page is part of the creative sequence. The measurement consequence is that a test built without traffic segmentation does not just fail to find an effect. It produces a false negative that licences neglect.
An illustrative teardown
Take a hypothetical midcore title running a six-week store listing experiment on Google Play with two screenshot sets. Control leads with a cinematic hero shot. Variant leads with a gameplay frame that matches the current best-performing video hook. The result arrives: variant up by a fraction of a percentage point, inside the confidence interval, declared no effect.
Break the same six weeks by acquisition channel and a different picture appears. Browse and search visitors converted almost identically on both sets, because the hero shot and the gameplay frame are both competent generic pages. Visitors from the video campaign whose hook the variant matched converted noticeably better on the variant. Visitors from a second, older video campaign with a different hook converted noticeably worse on it, because the variant page now looked like a different game from the one in their ad. The two paid effects roughly cancelled. The blended line stayed flat.
No amount of statistical care on the blended result would have rescued this. The experiment was answering "which page is better for everyone" when the only question with a real answer was "which page is better for whom".
A decision rule for the next test
Before launching a store page experiment, apply a routing test. If you cannot say which traffic source the variant is designed to serve, you are not ready to run it.
- If the variant is designed for browse visitors, run it as a default page experiment, but size it for browse volume alone, not total volume. That usually means a much longer run than the platform's estimate.
- If the variant is designed for a specific paid campaign, do not run it as a default page test at all. Route that campaign's traffic to a custom product page or a custom store listing and read conversion per campaign in the MMP or the platform console.
- If the variant is designed for search, stop. Search conversion is dominated by brand intent and the icon, and screenshot tests there mostly measure noise.
The general rule: default page tests measure the browse experience and should be judged on browse conversion. Everything paid should be tested through campaign-specific pages, with the page treated as a component of the creative rather than an independent variable.
What to read instead of the blended number
The single most useful figure a store page test can produce is conversion by page variant by campaign, read for the top three paid campaigns individually. If the platform's experiment tooling cannot provide that split, the fallback is to run the variant as a custom page assigned to one campaign and compare it against that same campaign on the default page over the same period, accepting that this is a quasi-experiment rather than a clean split.
The result to look for is not "variant wins" but "variant wins for campaign X and loses for campaign Y". That is the finding that changes how the team works, because it implies the store page should be built per concept, refreshed on the same cycle as the ads, and owned by whoever owns the creative rather than by whoever owns the listing. A team that has seen that split once stops asking whether the store page matters and starts asking why it was ever tested as if it had a single audience.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
measurement
media buying
·1 min read
Using predicted LTV in bids: disclosure checklist for the UA team
measurement
media buying
·1 min read
Blended ROAS targets that hide channel failure in F2P portfolios
measurement
media buying
·2 min read
Web-shop LTV with VAT-inclusive prices versus store net proceeds (labelled synthetic)
measurement
media buying
·1 min read
View-through attribution windows on F2P rewarded and interstitial traffic
More from the Measurement desk
measurement
·2 min read
Airbridge adds Amazon Ads as an app measurement channel
measurement
·2 min read
When to freeze a cohort for payback review (and when not to)
measurement
·1 min read
Web-shop purchaser quality vs store IAP purchaser quality
measurement
·1 min read