A/B test email so improvements compound
Guessing which subject line wins wastes sends; a clean test settles it.
AI can design the 4-week test calendar and compute the lift your list size can actually detect. You run it in Klaviyo and log results, because the real numbers are yours.
Fastest path: one prompt, end to end
🤖 AI prompt — paste into ChatGPT / Claude
You are a CRO analyst designing an email A/B testing program.
List size (typical single send): [#] Send frequency: [per week] Current open rate: [%]
Using MY numbers only:
1. Design a 4-week subject-line test calendar: ONE variable per test (week 1 length, week 2 personalization, week 3 emoji, week 4 curiosity vs clarity), plus the single metric each week is judged on.
2. Compute the minimum sample per arm to detect the lift I care about at 95% confidence and 80% power, from my list size only. State the baseline open rate you used and the lift you assumed. If my per-send list can't reach that sample in one send, say so and tell me exactly how many sends to pool.
3. Give a logging template with columns: test | variable | variant A/B | metric | result | winner.
Do not invent my open rate, do not declare a winner on tiny samples, and cite the assumption behind any sample-size number. If you can't compute it, tell me which input is missing.
Output: 4-week calendar + sample-size note + log template.
Or do it in 5 steps
- Change ONE variable per test. Subject length OR personalization OR emoji, never several at once, or you can't attribute the result to anything.
- Check the sample is big enough, honestly. A common practical floor is ~1,000 recipients per arm, but that only catches large gaps. At a 20% baseline, 95% confidence, 80% power, a ~10% relative lift (20% to 22%) needs ~6,500 per arm; a smaller 5% lift needs ~26,000; a 3% lift ~70,000 (suped.com, 2026). Use your ESP's significance indicator and don't over-trust small wins.
- Test the highest-leverage thing first: subject line (drives opens) before body tweaks (drive clicks). Fix the biggest lever first.
- Run a 4-week subject-line series, one test per week, so wins compound: week 1 length, week 2 personalization, week 3 emoji, week 4 curiosity vs clarity.
- Log every result in one sheet and roll winners into your default style. Record test, variable, both variants, the metric, and the winner. An untracked test teaches you nothing next quarter.
Worked example (labeled, from reader numbers only): detectable lift is a function of list size, not opinion.
At a 20% baseline open rate, 95% confidence, 80% power, a rigorous 2-point (10% relative) lift needs ~6,500 per arm, so ~13,000 per send.
A list of 8,000 split 50/50 (4,000/arm) can only reliably read a bigger gap in one send, or pool a 2-point test across ~3-4 sends.
A 2,000-person list (1,000/arm) hits the rough floor but not real significance, so pool it across several sends before you trust the winner.
Log every test; promote winners into your baseline monthly.