A/B testing on low traffic: what you actually can test
Most A/B testing advice assumes thousands of conversions a week. On a small store a classic split test almost never reaches significance, so its 'winners' are usually noise.
The move: test only big swings, run clean before/after periods with guardrails, and lean on qualitative methods that need five users, not five thousand.
Fastest path: one prompt, end to end
🤖 AI prompt — paste into ChatGPT / Claude
You are a CRO analyst for a low-traffic store. Use only my numbers; do not invent any.
My weekly numbers: sessions [#], conversions/orders [#], baseline conversion rate [%].
The change I want to test: [describe it].
Do this:
1. Estimate whether I have enough traffic to detect a realistic lift. Use my baseline rate and weekly conversions; state roughly how many conversions per variant a standard A/B test needs for a small vs large effect, and how many weeks that takes at my volume. Cite the rule of thumb you use.
2. Give a verdict: classic 50/50 A/B test, or not enough traffic.
3. If not enough, recommend the low-traffic alternative: (a) test only BIG swings (offer, hero message, price framing), never button colors; (b) a sequential before/after test with guardrails (same weekday mix, exclude promo spikes, minimum 2-3 weeks each side); (c) qualitative substitutes: 5-user tests, session recordings, on-site polls.
4. Warn me about peeking (checking daily and stopping at the first 'win') and how it manufactures false winners.
Do not guess my traffic. If a number is missing, ask.
Output: can-I-test verdict + the method to use + the specific guardrails for my test.
Or do it in 4 steps
- Do the sample-size reality check first. A store with 2,000 sessions and 40 orders a week runs at a 2% baseline. To detect a small lift (2% to 2.4%), a standard test needs thousands of conversions per side, or months of traffic. Knowing this stops you trusting a 'winner' that is just noise.
- Test only big swings. Low traffic detects only large effects, so test changes big enough to make one: the offer itself, the hero headline and image, price presentation (bundle vs unit, shipping in vs added). Skip button colors and micro-copy. Their true lift is too small for your volume to ever see.
- If you can't split, go sequential with guardrails. Run the old version 2-3 full weeks, then the new version 2-3 full weeks. Guardrails keep it honest: match the weekday mix, exclude promo or viral spikes, never compare a sale week to a normal one. It's weaker than a split test, but on low traffic it's often all you have.
- Substitute qualitative for volume. Five users doing a task out loud, ten session recordings, or a one-question poll show why people bail. They need five people, not five thousand. On a small store, watching real sessions beats waiting a year for significance.
Worked example (labeled): a store at 1,500 sessions and 30 orders a week (2% baseline) wants to test a new hero. A 50/50 split would take months to detect anything short of a huge lift, so a classic test is out.
Instead, run the current hero for 3 weeks (2.0%), then the new hero for 3 weeks (2.9%), with the same weekday mix and no promo weeks in either window. A ~45% relative jump on a big change is believable where a 5% blip would be noise.
Meanwhile five session recordings show shoppers missed the old value prop. That's qualitative evidence that agrees with the numbers.
On low traffic, only big swings are detectable; run clean before/after periods, never peek and stop early, and let five-user tests do the work volume can't.