Building an experimentation culture in a UA team
By Isaac Turner, Measurement Editor — Archive date: 5 min read
View author profile
Test registries, pre-registered hypotheses, sample size discipline and a decision log turn scattered UA tests into a real experimentation habit.
A UA team that changes a bid strategy, or tries a new creative concept, or shifts budget between channels is running a test in the loose sense: something changed and someone is watching what happens. Far fewer teams run experiments in the stricter sense. That version means writing the hypothesis down before the change goes live, agreeing a sample size in advance and recording the result whether or not it confirms what the team hoped for. The gap between those two modes is where a lot of UA budget quietly goes, spent on pattern matching rather than evidence.
It sounds academic until a team tries to answer one question honestly: what has this team learned in the last six months that it didn't already believe going in? Most teams can list campaigns they ran and channels they tried. Far fewer can point to a specific belief they held before a test and then dropped because the result contradicted it, and that is the practical symptom of an organisation testing without experimenting. It tends to travel with a UA budget that grows every quarter while the team's model of what works gets no sharper.
A test registry does the unglamorous work
The simplest structural fix is a test registry: a shared, append-only record of every test the team runs, what changed, what the team expected, what sample size or duration it agreed before launch, and what happened. No single entry is worth much. The value is in what the registry prevents over time, which is the same idea getting re-tested every few months because nobody remembers it already failed, and the selective memory that keeps only the wins visible in a team's collective sense of what works.
Pre-registering a hypothesis, even in a single sentence, forces a team to say what would count as a win before it sees any data. That closes off the most common form of self-deception in UA analysis: deciding after the fact that a middling result supports the change because some secondary metric moved the right way. A hypothesis written as "this creative concept will lift day-one retention by a meaningful margin over the current control" gives the team something it can be wrong about; a hypothesis written afterwards never does.
Sample size discipline and the decision log
Sample size discipline is the least exciting part of the process and the one teams skip most. Agree in advance how many installs, or how many days, a test needs before anyone reads the result, and the team stops itself from calling a test early on a number that only looks good because it hasn't yet had time to regress toward the account's real average. The threshold doesn't need to be a formal statistical power calculation, though larger teams increasingly run one. It does need writing down before the test starts, not negotiating after a promising early read.
A decision log sits alongside the registry and records something the registry can't: what the team decided to do with each result, and why. Suppose a test showed a small, statistically uncertain lift. Scaling it anyway might be perfectly reasonable if the downside was low and the concept fit a wider creative strategy, and the log captures that reasoning at the time, so six months later, when someone asks why a particular concept is still running, the answer exists instead of getting reconstructed from memory.
Rewarding the learning, not just the win
The cultural piece is harder to build than any of the process pieces above it. A team that only celebrates clear wins quietly teaches its analysts to avoid uncertain, high-information tests in favour of safe ones that will probably produce a modest positive. That's a bad trade. A test that cleanly disproves a popular internal theory, and saves the team from scaling something that would have wasted budget, is worth as much as a test that finds a winner; the team's incentive structure, however informal, needs to say so out loud rather than leave it implied.
The registry and the decision log, with the hypothesis written first, are the scaffolding. What sustains an experimentation culture past its first few months is whether a well-run test that failed gets treated in the following team meeting as differently valuable from a badly run test that happened to succeed, and whether the analyst who ran the failed test gets asked what they learned rather than why the numbers didn't move.
A managing editor running this discipline across a content operation rather than a UA desk ends up with almost the same structure: a registry of what the team tried, a stated expectation before publishing, a log of what the team decided once the result came in. The domain changes. The habit doesn't: write the hypothesis down before looking at the outcome, and treat a clean negative result as worth exactly as much as a clean positive one.
Related archive reading
These articles provide related context and remain subject to their stated review status.
Featured
Related posts
market intelligence
measurement
·1 min read
Retention percentile charts for F2P: what GameAnalytics-style cuts miss
market intelligence
measurement
·2 min read
Sensor Tower’s $82 billion figure is an IAP measure
market intelligence
measurement
·2 min read
The 52 billion download headline crosses platforms
market intelligence
measurement
·2 min read
One casual-gaming report contains several data populations
More from the Market Intelligence desk
market intelligence
media buying
·1 min read
Lunar New Year UA gates from an official calendar, not a CPI myth
market intelligence
media buying
·2 min read
WeChat Mini Games upgrades 2026 IAP / virtual-payment incentives for debut titles
market intelligence
media buying
·2 min read
WeChat Mini Games IAA incentives from 20 August 2026: 3-minute lifetime, 180-day option
creative strategy
market intelligence
·1 min read