How to pick a primary metric for a test
Choose the one number that decides your A/B test before it starts: tie it to the decision, check your tool reports it, add guardrails and write the rule down.
What this is for
Every test reports a dozen numbers, and some of them will favour each version by chance. The primary metric is the single number you agree, before the test starts, will decide the winner. Everything else is context.
Klaviyo's help centre puts it in one line when it asks you to set up a campaign test: pick the main metric you want to see improve. By the end of this guide you will have that metric, the rule for calling a winner and one or two numbers that must not get worse, all written down. source, checked 28 September 2026
Time needed and what you need
Time needed: about 30 minutes. Most of it is thinking through what you would actually do differently if the new version won.
- The change you plan to test and the hypothesis behind it
- The list of metrics your testing tool reports for this kind of test
- Your current figures for the candidate metrics, from Shopify's reports or your email platform
- Somewhere to write the test plan that you cannot quietly edit later, such as a shared doc with its history on
The steps
Do this before you build the test. A metric chosen after the results are in is a metric chosen to agree with you.
1. Start from the decision, not the dashboard
Finish this sentence: if the new version wins, we will... Roll out a new product page layout, switch the free delivery threshold, change how subject lines are written. The primary metric is whatever proves that decision was right. If you cannot name the decision, the test is not ready.
2. List the candidate metrics
Write down every number the change could plausibly move, from the step closest to the change to the step closest to money. For a product page change that might be add to cart rate, reached checkout rate, conversion rate and revenue per session. For an email it might be open rate, click rate and placed order rate.
3. Pick the one closest to money that the test can still measure
Metrics near the change react strongly and need fewer visitors; metrics near revenue matter more but are noisier and need far more traffic to show a real difference. Choose the metric closest to revenue that your traffic can measure in a few weeks. A test on a button's wording can be judged on add to cart rate. A test on price or a delivery threshold has to be judged on money, because it deliberately trades orders for order size.
- Changes the page: add to cart rate or conversion rate
- Changes the basket: revenue per session
- Changes the checkout: checkout conversion rate
- Changes an email's content: click rate or placed order rate
4. Check your tool actually reports it
A metric you cannot read at the end is no use. Shopify's rollout experiments report a fixed set: for a theme change, conversion rate, bounce rate, reached checkout rate and add to cart rate; for a checkout change, checkout conversion rate. Revenue is not on either list, so a revenue question needs a tool that reports it. source, checked 28 September 2026
Email tools have their own lists. Klaviyo decides campaign tests on one of three. source, checked 28 September 2026
5. Be wary of open rate
Apple Mail Privacy Protection loads email images automatically, so opens are over-counted. Klaviyo suggests a separate report if a large share of your opens come from Apple Mail. Use open rate only for tests that can only affect the open, like a subject line, and prefer click rate or placed order rate for anything else. source, checked 28 September 2026
6. Add one or two guardrail metrics
A guardrail is a number that must not get meaningfully worse, even if the primary metric improves. A bigger discount banner can raise conversion rate while cutting average order value; a more urgent email can raise clicks and unsubscribes together. Pick one or two guardrails, say how much of a drop you would accept, and do not use them to pick the winner.
7. Set the smallest change worth acting on
Decide the minimum detectable effect: the smallest lift on your primary metric that would be worth the effort of rolling the change out. It sets how many visitors the test needs. You can find that number with the A/B test significance calculator before you start.
8. Write the rule down and keep to it
One sentence: version B wins if its revenue per session is higher at the planned end date, the difference passes the significance check, and conversion rate has not fallen by more than the limit you set. Save it with the date. When the results arrive, judge the test by that sentence, even if another metric looks more exciting. Then go back to run the A/B test on Shopify.
A worked example
For a home fragrance store raising its free delivery threshold from $50 to $65
Worked example (an invented store, with invented numbers)
The owner expects fewer orders but bigger baskets, so she writes down revenue per session as the primary metric and conversion rate as a guardrail, allowed to fall by no more than 10% of its current level. Her testing app reports orders, sales and sessions for each version, which a rollout experiment would not.
After two weeks each version has 10,000 sessions. The control took 250 orders at an average of $60, which is $15,000 in sales or $1.50 per session. The new threshold took 230 orders at an average of $72, which is $16,560 or about $1.66 per session.
Judged on conversion rate, the new threshold lost: 2.5% down to 2.3%, a relative drop of 8%. Judged on the metric she chose, it won by about 10.4% per session. She also checks that the lead held in both weeks: $1.49 against $1.64 in week one and $1.51 against $1.67 in week two. The 8% conversion drop is inside her 10% limit, so she keeps the $65 threshold.
Common mistakes
Choosing the metric after the results are in. With enough numbers on a report one of them will favour the new version by chance, and picking it then is how false wins get shipped.
Using clicks or add to carts to judge a change that is really about money, such as price or a discount. Those metrics can rise while revenue falls.
Picking a metric your traffic cannot move in time. If revenue per session would need three months to show a real difference, choose a metric nearer the change, or test a bolder change.
Testing several things in one email and then arguing over which metric counts. Klaviyo's own advice is to test one variable at a time, which also keeps the choice of metric simple. source, checked 28 September 2026
Doing this with Tilly
An A/B test built by Tilly's experiments agent costs 10 credits, and the page change it builds lands on a preview copy of your theme for you to approve before anything is live. Choosing the primary metric and the rule for a winner is still a decision you make and write down first.
Questions people ask
Can a test have two primary metrics?
It should not. With two, you will be tempted to call whichever one won, which doubles your chances of acting on a fluke. Keep one primary metric and treat the rest as guardrails or background.
Which metric should an email subject line test use?
Open rate, because the subject line is only seen before the open. Klaviyo recommends open rate for subject lines and sender names, and click rate for content changes such as a button. source, checked 28 September 2026
Is bounce rate a good primary metric?
Rarely. It is easy to move and it is reported in Shopify's theme experiments, but fewer one-page visits does not mean more orders. Use it as a guardrail on a landing page change, not as the number that decides the test.
See this on your own store
Paste your store URL. The first pass takes about a minute and needs no account. Save a card (nothing is charged) and Tilly reads the whole store and works out who buys from you.