Most tests are stopped for one of two reasons: someone got impatient, or someone got excited. Neither is a statistical event. This piece is about the third reason — the one that holds up when the result reaches your P&L.
The question sounds like a scheduling problem and it is really a measurement problem. Two teams running the same test on the same page, with the same traffic, can reach opposite conclusions purely by choosing different days to look. What follows is how to stop that from happening to you.
Why duration matters more than volume
The instinct is to think of a test as a bucket that fills with visitors, and to assume that once the bucket is full the answer is inside it. That model is wrong in a way that costs real money, because it treats every visitor as interchangeable.
They are not. A Tuesday morning visitor arriving from a branded search is a different person, with a different intent, from a Saturday evening visitor arriving from a cold interest-targeted ad. If your test has only seen one of them, it has an answer to a question you did not ask.
Sample size, and what it does not tell you
Every calculator you have used asks for the same three inputs and returns a number of visitors per variation:
- a baseline conversion rate, taken from a period you hope was representative;
- a minimum detectable effect, which is a business decision dressed as a statistical one;
- a confidence level, almost always left at the default.
That number is necessary and it is not sufficient. What the calculator assumes is that your visitors arrive as a random sample from a stable population. In paid media that assumption is close to never true. Campaign budgets shift, creative fatigues, audiences saturate, and the mix of traffic reaching your page on day two is not the mix reaching it on day nine.
A
The two-cycle floor
A single cycle tells you what happened. Two cycles tell you whether it happens again. Below two complete cycles you are not measuring an effect, you are measuring a week.
Novelty and the shape of a false winner
Returning visitors react to change as change. A new headline gets attention because it is new, and that attention is indistinguishable, in the first few days, from the attention a genuinely better headline would earn.
The shape is recognisable once you have seen it: a strong early lift that decays steadily as the proportion of visitors who have already seen the variation grows. If you stop the test during the decay, you record the average of a spike and a fall, and you ship it.
Confidence bands narrow as a test accumulates cycles. The point at which they stop moving - not the point at which they first look convincing - is when a result can be believed.Source: Visiopt test archive, 56 documented tests · Chart: Visiopt Research
When to stop
The useful question is not “has it reached significance?” but “has it stopped moving?” A result you can act on has three properties, and significance is only one of them:
- It has been stable across two complete cycles, not just statistically significant at one moment.
- It holds in the segments that matter, not only in the aggregate.
- It did not depend on the day you happened to look.
A result that reached 95% on day four and has been drifting since meets none of them.
What early stopping actually costs
The cost is not the lost test. The cost is everything downstream of a wrong decision: the traffic you send to a losing page while you believe it is winning, the follow-up tests you build on a false baseline, and the confidence you spend on a number that was never there.
Final takeaway
A test does not become believable when it becomes significant. It becomes believable when it stops changing its mind. Significance is a threshold; stability is evidence. If you only have room to remember one of them, remember the second.
| Factor | What to look for | Example |
|---|---|---|
| Headline | Is the value clear? | “Increase your conversions” |
| CTA | Is the next step obvious? | “Start free trial” |
| Form | Is it asking too much? | 3 fields instead of 8 |
If you want to learn more about conversion optimization, see our conversion optimization guide.
