A/B testing and experiments
Create and run A/B tests natively in Ordinary — variants, traffic split, page and audience targeting — scored on your real reconciled revenue with peek-safe statistics.
A/B testing and experiments
Run A/B tests directly in Ordinary and score every variant on the real orders your store actually shipped — not clicks, not sessions, not a vanity uptick. Set a test up in a couple of minutes; Ordinary handles the traffic split, the statistics, and the revenue attribution.
Left nav → Experiments.
Create a test in Ordinary
Click New experiment, give the test a name and the page it runs on, and it’s created as a draft that opens straight onto its test page. (Splitting visitors between two different URLs, scoping to a channel up front, or importing from a legacy tool? The create dialog links to a full form that covers those in one pass.)
Everything about a test — reporting and configuration — lives on one page, organised into tabs. Results shows how the test is doing (it’s where you land once a test is running); the config tabs below are editable while the test is a draft or paused:
-
Test groups — add 2–10 groups and set the traffic split (50/50, 80/20, whatever you need). The first group is your control — your current page, unchanged. Drag the split bar to move traffic between groups in 5% steps, or type exact percentages. You can also change the split while the test is running — ramp a winner up gradually, or set a group to 0% to stop sending it traffic without ending the test. Shoppers already in the test keep their group, so only new visitors are affected, and each split is scored from the moment you save it. Which groups exist and where they point stay fixed once a test starts.
-
Modifications — what each group changes on the page. Author edits visually (click any element on your storefront and change its text, style, image, or visibility), or ship variant content in your theme code using the attribute strings from the results page — both are listed here per group.
-
Targeting — who enters the test and where it runs. Limit entry to a device type (mobile, desktop, tablet — e.g. test a mobile-only layout change on mobile traffic alone), a visitor type (new or returning — e.g. run a first-time-buyer offer at people who’ve never visited, without touching your repeat customers), and/or a traffic source, and the currency they’re shopping in. Sources are grouped: Paid (Meta ads, Google ads, TikTok ads, Microsoft ads, other paid social — Pinterest, Snapchat, X, LinkedIn, Reddit), Owned (Klaviyo, other email/SMS), and Earned (organic search, organic social, AI assistant, referral, direct). So you can run a landing-page test on TikTok ads only, or try different copy on people arriving from ChatGPT. A shopper counts as returning once they’ve visited your store in an earlier session; browsing privately, on a new device, or with site data cleared reads as new. Everyone else sees your normal site; the filters decide who enters, and once a shopper is in they keep their assigned group on every return visit, whatever device or source they come back through — so a new-visitor-only test keeps scoring the people it enrolled after they become returning visitors. Set the page URL pattern — a trailing
*matches a whole section (e.g./products/*). For multi-page tests this tab also holds persistence: extra pages where shoppers already assigned to a treatment group are taken back to their group’s page (new visitors are never pulled in from these pages, control shoppers are never redirected, and the cart and checkout are never touched), plus a sticky window — how many days each shopper keeps their assigned group, counted from their first visit (default 365).Currency targeting is the one to reach for on a price test. It matches the currency the shopper is actually seeing prices in, so someone in Ireland checking out in GBP lands in your GBP group — which is the grouping a pricing question is really about. Shoppers in a currency you haven’t selected see your normal page.
“Also run on these pages” lets one test start on more than one page. A test has a single page URL, which covers “this page” or “this section” — but an offer test usually spans a product page and the landing page that sells it, and those don’t share a prefix. Add the extra paths here, one per line; they follow the same rules as the page URL, so a trailing
*still matches a whole section. If you find yourself listing many pages, a*pattern is usually the better tool.This tab also holds the exclusion group. If two of your tests can reach the same shoppers, a shopper who lands in both makes each result harder to trust — you can’t tell which change moved the number. Give both tests the same exclusion group label and no shopper will ever be in both: each one is put into one test or the other, and stays there. Shoppers already in a test keep it, so adding a label to a test that’s already running never removes anyone — it only affects who joins from that point on. Leave it empty (the default) when your tests don’t overlap, or when you’re happy for them to run together. Each test then reaches its full audience and reads its result sooner.
-
Goals — the primary metric that decides the winner (revenue per visitor by default, or conversion rate / average order value), optional guardrail metrics that must not regress (refund rate, bounce rate, add-to-cart rate, 30-day repeat purchase), your hypothesis, and the attribution window.
The attribution window is how long after a shopper first sees the test their orders still count toward it. The default is 14 days, and it fits almost every test — most orders happen within a day or two of the visit. Lengthen it only if your shoppers genuinely take longer to decide, and know the trade: a long window also starts counting repeat purchases that would have happened anyway, which flatters whichever group happens to contain more of your existing customers.
It’s easy to mix this up with the sticky window on the Targeting tab, because both are counted in days from the same moment. They do different jobs: the sticky window controls how long a shopper keeps seeing their assigned version, and the attribution window controls how long their orders count. Ordinary won’t let you set the sticky window shorter than the attribution window, since that would mean showing someone the original page while their purchases still counted toward the version they used to see.
-
Preview & QA — per-group links that force a page load into that group so you can QA it. Preview visits write no cookie and record no exposure, so your test data stays clean. Links go live once the test starts.
-
Theme changes — if this test changes wording or images in your theme, this shows you exactly what Ordinary added, line by line, and lets you take it back out. See Removing a test’s changes from your theme.
Click Start when you’re ready. Ordinary serves the split through your store — no separate testing tool, no DNS or tag-manager setup. Bucketing is sticky (a shopper always sees the same group), and results start flowing within minutes.
One thing to know when you end a test: shoppers go back to your normal page immediately, but orders from people who saw the test can still land for the length of the attribution window. So the numbers keep settling for up to 14 days (or whatever window you set) after you stop it. That’s deliberate — cutting it off at the moment you press Stop would throw away real orders from shoppers who saw the test just before it ended. If a result is close, wait out the window before calling it.
How many tests you can run
The Free plan includes one A/B test. It’s there so you can run a real test on your own store, with your own traffic, and see what the feature does before paying for it. Ending or archiving that test doesn’t free up another one.
Starter and Advanced remove the limit — run as many as you like, at the same time or one after another.
If you’re on Free and the New experiment button is greyed out, you’ve used your test. Your test and its results stay readable whatever plan you’re on.
What the dashboard shows
The Experiments list shows every test, with a headline metric and a plain-English verdict. Clicking into a test opens the detail view:
- Per-arm results for the visitors who saw each variant — orders, conversion rate, average order value, revenue per visitor.
- Lift — how the variant performs against the control, with a plain-English “range of likely results” (the statistical confidence interval, named so it’s actually readable).
- Peek-safe significance — a sequential-testing method that lets you check results early without inflating the false-positive rate the way classic significance tests do.
- Faster conclusions — for revenue tests with enough data, Ordinary applies a variance-reduction technique (CUPED) that can shorten the time to call a winner by 20-30%. The display just shows tighter numbers; the math is in the About the math expander.
- Allocation check — if your split looks suspicious (you set 50/50 but actual exposure is 65/35), the dashboard flags it so you don’t trust biased results.
- Funnel by arm — sessions → add-to-cart → checkout → orders, broken out per variant, with a step-to-step table showing what share of each step advanced and the gap between arms — so you can see WHERE a test wins or loses (a cart-adds win that dies at checkout is a different decision than a checkout-completion win).
- Guardrails — the metrics you chose to watch for damage (refund rate, bounce rate, add-to-cart rate, 30-day repeat purchase, revenue per visitor). Each shows both arms, the difference, and a status that understands direction — a refund-rate increase is flagged, not a decrease. Metrics that can’t be computed honestly yet (for example repeat purchase before any visitors are 30 days old) say “No data yet” instead of pretending everything is fine.
- Value breakdown — where a revenue move came from: average order value, units per order, product and shipping revenue per order, discounting, orders refunded so far, subscription share, and repeat purchases inside the window. Diagnostic by design — the winner call stays with the headline lift.
- Profit per visitor — a per-arm contribution P&L: product revenue and cost of goods (from the costs already on your catalog — no separate COGS entry step), shipping revenue and cost, fulfillment and transaction fees, and contribution profit per visitor. If cost coverage on an arm is below 70% of revenue, Ordinary shows the coverage figure and how to raise it instead of a misleading number.
- Ad spend context — what the store spent on Meta and Google while the test ran, each arm’s paid-traffic share, new customers, and effective cost per new customer, plus what the observed lift range would mean per month at your recent revenue pace — spend and test results on one surface.
- Revenue per visitor over time — a line per variant; drift can hint at allocation problems or seasonality.
Every number uses the same reconciled revenue and attribution as the rest of your reports — every order counted is a real, paid Shopify order.
How to interpret “Too early”
Ordinary won’t pretend a test is significant when it isn’t. The dashboard shows “Too early” until at least 1,000 visitors have been exposed and the statistical confidence threshold is met. There’s no way to dismiss this — the math is the math.
To plan a test before running it, use the public sample size calculator: plug in your baseline conversion rate and the smallest lift you’d want to detect, and it tells you how many visitors per variant you need.
Settings and defaults
Open Settings → Experiments for org-wide preferences:
- Default primary metric — usually revenue per visitor; some teams prefer conversion rate or AOV.
- Default guardrails — secondary metrics checked alongside the primary so you catch trade-offs (a variant might “win” on conversion rate but lose on refund rate, which is a wash).
- Conversion event — by default Ordinary counts
checkout_completed. Switch tocheckout_startedoradd_to_cartfor upper-funnel tests. - Bridge attribute names — only relevant if a third-party theme test uses non-standard attribute names.
You can also override defaults per-test from the About this experiment card on the test detail page.
How content changes reach your storefront
Tests take one of two routes to your live store, depending on what the variant changes.
Styling and visibility changes — colour, size, spacing, hiding something — are applied as the page loads. Nothing in your theme changes.
Content changes — a different headline, a different image, a changed link — are made in your theme. Every one of those edits:
- Only adds. Your original wording stays exactly as you wrote it. Ordinary puts the variant alongside it, so a shopper sees one or the other. Nothing of yours is replaced or deleted. If a change can’t be made that way, Ordinary won’t make it.
- Is backed up first. Ordinary saves your theme exactly as it was before changing anything, so it can be put back exactly.
- Is checked, and undone if anything’s wrong. Ordinary checks your page after the edit. If it doesn’t come out right, the change is removed and the test doesn’t start.
Everything Ordinary adds is labelled as ours and kept together, so it’s easy to spot in the theme editor.
Seeing exactly what changed
Open the test and go to Theme changes. It lists every file Ordinary touched for that test and shows the change line by line, the way a developer would review any other edit — added lines marked, your original lines alongside them. Each file shows when the change was made and whether it’s still in place.
If a test has never changed your theme, the tab says so.
Removing a test’s changes from your theme
Ordinary does not remove them on its own — not when you end a test, and not when you uninstall. They stay in your theme until you take them out, so that ending a test never quietly edits your files.
To take them out: open the test, go to Theme changes, and click Remove from theme. Each file goes back exactly as it was before the change.
What stays behind until then does nothing. Once a test isn’t running, the alternative versions stay hidden and every shopper sees your original content. There’s nothing broken in the meantime — it’s your own wording on the page either way.
If you’ve edited the file yourself since Ordinary changed it, Ordinary won’t overwrite your work. It stops and tells you it couldn’t put that file back, rather than risk removing something you wrote. Get in touch and we’ll help you sort it out.
Before you uninstall: an uninstalled app can no longer edit your store, so remove the changes first if you want them gone. If you’ve already uninstalled, reinstall and remove them, or contact support and we’ll walk you through taking the labelled blocks out by hand.
Which parts of your theme change
Only your live theme, and only the section your tested headline, image or link sits in. Ordinary never touches your theme settings, your templates, or your uploaded files. If it can’t tell for certain which part of your theme creates what you picked, it won’t start the test.
Data coverage and cookie consent
Experiment data is captured via the same pixel events that power the rest of Ordinary’s analytics. If your storefront uses a cookie banner integrated with Shopify’s Customer Privacy framework — common for EU/UK/EEA-facing stores — shoppers who decline cookies generate no pixel events, so they’re invisible to the experiments platform.
In practice:
- Sample sizes for tests with strict-region traffic will be lower than your overall visitor count suggests. Plan accordingly.
- The shoppers in your results are the accept-cookies subset. People who decline might behave differently — a technical selection bias, but every analytics platform on Shopify has the same limitation, and the framework is what Shopify requires for compliance.
- Non-strict-region traffic (US, Canada, Australia, most of Asia and South America outside specific markets) sees no banner by default and produces a full event stream.
If a test needs a specific sample size, factor in your accept-cookies rate. A test with 40% EU traffic and a 60% accept rate captures roughly 76% of your visitors — plan for ~30% more traffic, or run it ~30% longer.
See How Ordinary handles your visitor and customer data for the full posture.
Split your other reports by variant
While a test runs, Ordinary can split your existing reports by variant. Look for the Experiment dropdown in the page actions on Orders, Customers, and the GMV chart — selecting a test breaks every row out by which variant the visitor saw. Handy for spot-checking that the per-variant aggregates agree with live order data.
Frequently asked
Q: My test isn’t collecting data. Why not?
A: If it’s a test you built in Ordinary, make sure you clicked Start (drafts don’t run) and that real traffic is hitting the URL pattern you set — give it ~10 minutes after the first shopper lands. If a test is running but the Results tab stays empty for longer than that, contact support with the test name and we’ll trace it.
Q: Can I change the traffic split while a test is running?
A: Yes — that’s how you ramp a rollout. Move the split on the Test groups tab and save; shoppers already in the test keep their group, and only new visitors are divided by the new percentages. The allocation check scores each split separately from the moment you save it, so a deliberate ramp doesn’t read as a traffic-split problem. Setting a group to 0% stops new traffic reaching it without ending the test. What stays fixed once a test starts: which groups exist, and where they point.
Q: I changed my test variants mid-flight. What happens?
A: Changing the content of a variant (as opposed to its traffic share) is a different matter. Ordinary flags this as an assignment conflict on the affected visitor records and surfaces a warning if the conflict rate is non-zero. Best practice is to end the original test and start a fresh one rather than reshuffling buckets — but Ordinary won’t quietly mix the data.
Q: How long does Ordinary count an order as belonging to a test?
A: 14 days from when a shopper first saw a variant. A purchase 18 days after exposure falls outside the window and isn’t attributed to the test. 14 days is the industry-standard experiment attribution window.
Q: Can I export per-visitor exposure data?
A: Yes — admins can open the Raw exposure log from each test’s detail page. The export includes the visitor’s anonymous identifier, the variant they saw, when they first saw it, and which event triggered the exposure.