How to know when an A/B test is done
An A/B test is done when it reaches the sample and full weeks you planned before it started, not when one side first looks ahead. How to set that line and read it.
What this is for
Most bad test decisions come from stopping at the wrong moment. Stop too early and you ship a fluke; let it run forever and half your visitors keep seeing the losing version. Google's own guidance for site owners names the two things that set the right length. source, checked 28 September 2026
This guide turns that into a rule you set before the test starts: a sample size, a whole number of weeks and a significance check at the end. When all three are met, the test is done, whatever the result.
Time needed and what you need
Time needed: about 45 minutes, split between the start and the end of the test. Half an hour to set the stopping rule before launch, and a quarter of an hour to check the result once the planned end date arrives.
- Your baseline rate for the primary metric, from at least the last few weeks
- The smallest lift you would act on, decided in advance
- Your average daily visitors to the page or email being tested
- The free A/B test significance calculator, or the one built into your testing tool
The steps
Steps 1 to 3 happen before the test starts. Steps 4 to 9 happen while it runs and when it ends.
1. Write the stopping rule before you launch
Decide four things and write them down with the date: the primary metric, the smallest lift worth acting on, the confidence level, and the number of visitors each version needs. This is called a fixed-horizon test: you decide in advance when you will look, and you only make the call then.
2. Work out the sample size
The sample size depends on two numbers: your baseline rate and the minimum detectable effect. The lower your baseline and the smaller the lift you want to catch, the more visitors you need. To get the figure, enter a trial pair in the A/B test significance calculator: your baseline for version A, the rate you would act on for version B. It shows how many visitors each version needs, at the usual confidence and power.
3. Turn it into whole weeks
Divide the sample per version by the visitors each version gets per day, then round up to whole weeks. Shoppers behave differently on a Sunday evening than on a Tuesday morning, and a test that covers Friday to Monday once but Tuesday to Thursday twice is lopsided. Google Ads gives its advertisers a similar rule of thumb for how long to let an experiment run. source, checked 28 September 2026
4. Do not stop the first time it looks significant
Checking every day and stopping the moment the calculator says significant is called peeking, and it is the most common way tests go wrong. Early in a test the numbers swing widely, and if you check often enough one of those swings will cross the line by chance. You can watch the test to make sure nothing is broken. Just do not act on the result until the planned end.
5. Check that the test ran cleanly
At the end, before you look at the winner, check the plumbing. Did each version get roughly the share of visitors you set? Did anyone edit either version during the test? Was there a site outage, a stock-out or a promotion that only one version showed? If the split is badly uneven or something changed midway, fix it and run the test again rather than trusting the result.
6. Run the significance check once, at the planned end
Enter each version's visitors and conversions in the calculator at the confidence level you chose. A p-value below the line means a difference this large would be unlikely if the two versions really performed the same. Google Ads, for comparison, tests its experiments in both directions at the 95% level. source, checked 28 September 2026
7. Read the range, not only the verdict
The calculator also gives a range for the lift. A result that is significant but whose range runs from a tiny gain to a large one tells you the change helps, not by how much. If the low end of the range is too small to be worth the change, you have a win on paper and little in practice.
8. Make the call and move on
If the new version won, ship it. If it lost, roll it back. If there was no clear difference, keep whichever version is simpler or cheaper to run and test something bolder next. Klaviyo gives similar advice for email tests that come back not significant: retest a couple of times, then move to a different question. source, checked 28 September 2026
9. Remove the test from the site
Once you have decided, take the losing version, any variant URLs and any testing scripts off the store. Google asks for this, and a forgotten test keeps splitting your visitors for no reason. source, checked 28 September 2026
A worked example
For a jewellery store with 2,400 product page sessions a day
Worked example (an invented store, with invented numbers)
The store tests a new product page layout. Its baseline conversion rate is 2.5%, and the owner decides a relative lift smaller than 20% (2.5% to 3.0%) is not worth acting on. The calculator says each version needs 16,792 visitors. At a 50/50 split each version gets 1,200 a day, so 16,792 ÷ 1,200 = 13.99, which rounds up to 14 days: exactly two full weeks.
On day 5 the owner peeks. Version A has 138 orders from 6,000 visitors (2.3%) and version B has 174 from 6,000 (2.9%). The calculator says p = 0.039, below the 0.05 line, but warns that each version has far fewer visitors than a lift this size needs, and the lift range runs from +1.3% to +50.8%. The owner leaves the test alone.
On day 14 each version has 16,800 visitors. A has 420 orders (2.5%) and B has 504 (3.0%), a relative lift of 0.5 ÷ 2.5 = 20%. The calculator gives p = 0.0051 and a lift range of +6.0% to +34.0%. The planned sample, the two full weeks and the significance check are all met, so the test is done and B ships.
Common mistakes
Stopping at the first significant result. If you check daily and stop on the first win, you will ship far more flukes than your confidence level suggests. Decide the end date first and keep to it.
Running part weeks. Four days of a test that starts on a Friday is mostly weekend traffic. Always run whole weeks, even if the sample is reached on day 10.
Treating a probability or promising label as a win. Klaviyo, for example, only labels a campaign test statistically significant once each version has reached a minimum number of recipients and the win probability is high; between 75 and 89 percent it calls the result promising and suggests running the test again. source, checked 28 September 2026
Letting a test run on and on because the result is not what you hoped. If the planned sample is reached and there is no clear difference, that is the answer. Google also warns that an experiment left running for an unnecessarily long time can look like an attempt to mislead search engines. source, checked 28 September 2026
Questions people ask
What if my store never has enough traffic to finish a test?
Then test bigger changes, which need fewer visitors to detect, or test on your busiest page. If even that would take months, make the change based on customer feedback and your product page audit, and watch the trend rather than pretending a small test proved it.
When is an email A/B test in Klaviyo done?
Klaviyo calls a campaign test statistically significant when each variation has reached at least 50 recipients and the win probability is at least 90%. Below that it labels the result promising, not significant or inconclusive. source, checked 28 September 2026
What does statistically significant actually mean?
That a difference as large as the one you saw would be unlikely if the two versions performed the same. Google Ads puts it simply for its own experiments: the result is probably not chance and should hold if you roll it out. source, checked 28 September 2026
See this on your own store
Paste your store URL. The first pass takes about a minute and needs no account. Save a card (nothing is charged) and Tilly reads the whole store and works out who buys from you.