Most email A/B tests fail because the sample is too small or the metric is wrong. Test one thing at a time, on a sample large enough to detect a meaningful difference, and judge on the metric that matches the variable. Subject lines move opens; creative and CTAs move clicks and revenue.
What is worth testing
- Subject line and preview text (judge by open rate).
- From name (judge by open rate).
- Headline and hero (judge by click rate).
- Offer and CTA (judge by click and conversion).
- Send-time (judge by open and click for that send).
Sample size, roughly
To detect a one-percentage-point lift in open rate at typical baselines you usually need tens of thousands per variant. For click and conversion lifts, more. If your list is small, run more tests over time and look at the trend, not a single send.
How to call a winner
- Pre-commit to the sample size and the metric.
- Run the test. Do not peek and stop early.
- Pick the winner only if it beats both the control and the noise.
- Send the winner to the remainder of the list.
- Log the result so you build a real testing history.
Common traps
- Testing two variables at once. You learn nothing.
- Running a winner-take-all on 500 contacts. That is noise.
- Ignoring revenue per email and chasing open rate forever.
Frequently asked questions
How many variants should I run?
Two or three. More splits the sample too thin.
Can AI choose the winner?
AI can suggest, but the call should be based on a pre-committed sample and metric, not a vibe.
How often should I test?
On most campaigns. The cost is small and the compounding learning is large.

