All articlesMetrics & Benchmarks

    How to A/B Test Emails Properly

    Most email A/B tests are not tests — they are noise read as signal. This guide covers writing a real hypothesis, choosing variables by impact, splitting the list correctly, the sample sizes needed to detect a given difference, why open rate is the wrong success metric, and how to build a testing program instead of a testing habit.

    How to A/B Test Emails Properly
    Erin Moore
    Erin Moore
    September 18, 20269 min read
    Share:

    A/B testing emails means splitting your audience randomly, sending each half a version that differs in exactly one element, and measuring which produces more of the outcome you care about. Testing properly requires a hypothesis, one variable, a large enough sample to detect a real difference, and a success metric further down the funnel than opens.

    Most email A/B tests are not tests

    Here is the pattern that plays out in thousands of accounts: someone sends subject line A to 500 people and subject line B to 500 people. A gets 22.4% opens, B gets 20.1%. A "wins." The team adopts a rule about emoji usage and moves on.

    That result is noise. With 500 recipients per arm and open rates around 21%, a 2.3-point gap is well inside the range you would get by splitting one identical email in half. The team just built a house style on a coin flip.

    Proper testing is not harder — it is just more disciplined about four things: what you test, how you split, how long you wait, and what you count as winning.

    Step 1: Write a hypothesis, not a curiosity

    A hypothesis has a mechanism and a prediction. "I wonder if a shorter subject line does better" is a curiosity. "Because most of our list reads on mobile and our subject lines truncate at 40 characters, a version under 40 characters will produce more opens" is a hypothesis.

    The difference matters because a hypothesis tells you what to do with a negative result. If the short version loses, you have learned something about your audience. If you were only curious, you have learned nothing and you will run the same test again in six months.

    Step 2: Test one variable, and pick a variable worth testing

    Change exactly one element. Two changes give you a result with no explanation. And prioritize by impact — some variables move outcomes far more than others.

    VariablePrimarily affectsTypical impactTest priority
    Offer / core value propositionClicks, conversionsLarge1
    Segment / audience targetingEverythingLarge1
    Email structure (long vs. short, single vs. multi-CTA)ClicksMedium-large2
    Subject line approachOpensMedium2
    Send day / timeOpens, clicksMedium3
    From nameOpensMedium3
    CTA copy and placementClicksSmall-medium3
    Images vs. text-onlyClicks, deliverabilitySmall-medium4
    Button colorClicksVery small5

    Notice where button color sits. Teams test it constantly because it is easy, and it almost never produces a detectable effect at real-world list sizes. Spend your test budget on offer and segmentation.

    Step 3: Split the list correctly

    • Randomize, do not slice. Splitting by signup date, alphabetically, or by contact ID introduces systematic differences between arms. Your platform's built-in A/B testing handles randomization for you.
    • Decide the split shape up front. A 50/50 split across the whole list gives maximum statistical power. A 20/20/60 split — test on 40%, send the winner to the remaining 60% — gives you a business outcome but less certainty.
    • Send both arms simultaneously. If A goes out at 9am and B at 2pm, you have tested send time, not your variable.
    • Hold segments constant. Do not run a test where one arm skews toward recent subscribers.
    • Do not peek and stop early. Checking results hourly and calling it when you see a gap is the fastest way to manufacture false winners.

    Step 4: Get the sample size right

    This is where most tests fail before they start. The smaller the effect you want to detect, the more recipients you need. A rough guide for detecting a difference in open rate, assuming a baseline around 25%, at 95% confidence (α = 0.05) with 80% power — the conventional defaults. Tighten either and the numbers climb:

    Difference you want to detectApprox. recipients needed per armRealistic for
    10 percentage points (25% → 35%)~350Almost any list
    5 percentage points (25% → 30%)~1,300Mid-size lists
    2 percentage points (25% → 27%)~7,500Large lists only
    1 percentage point (25% → 26%)~29,000Enterprise volume

    These are approximations for a two-sided comparison at conventional confidence levels, and click-rate tests need substantially more volume because the baseline is lower — a 3% click rate needs several times the sample of a 25% open rate to detect the same relative change.

    The practical implication: if you have 3,000 contacts, stop trying to detect small differences. Test big, structural changes where a real effect would be large enough to see. Small-list testing is not impossible — it just means testing bolder ideas.

    Step 5: Measure the right outcome

    Open rate is the most-used and least-reliable email metric. Privacy features that pre-fetch images inflate opens for a meaningful share of recipients, and that inflation is not evenly distributed across your list. A subject line test judged on opens alone can show a winner that produces fewer clicks.

    Use this hierarchy:

    1. Revenue or conversions per recipient — the real answer when you can attribute it.
    2. Unique click rate — the best available proxy for genuine interest.
    3. Click-to-open rate — useful for isolating body copy from subject line effects, with the same open-inflation caveat.
    4. Unsubscribe and complaint rate — always check. A variant that wins on clicks and doubles unsubscribes is not a winner.
    5. Open rate — directional only, and only for subject line and from-name tests.

    Wait at least 24–48 hours before reading results. Email engagement has a long tail and the ranking of two variants frequently flips between hour four and hour thirty.

    For subject line tests specifically, screen your candidates before they consume a test slot. Running two lines through the subject line tester catches length and spam-trigger problems, so your test compares two viable ideas rather than one good line against one that was never going to be delivered.

    Step 6: Build a testing program, not a habit of tests

    A single test is worth very little. A sequence of tests that build on each other is worth a lot. Structure it:

    1. Keep one log. Date, hypothesis, variable, arm sizes, metric, result, and what you concluded. A spreadsheet is fine.
    2. Test the same hypothesis across three or four campaigns before adopting a rule. Consistent direction across sends is stronger evidence than one significant result.
    3. Re-test your rules annually. Audiences change, clients change, and a truth from 2024 may be false now.
    4. Record the losers. The list of things that did not work is more valuable than the list of things that did, because it stops you repeating them.
    5. Separate exploration from exploitation. Roughly 80% of sends should use what you know works; 20% should test something genuinely new.

    Common mistakes, in the order they cost the most

    • Declaring winners on samples too small to detect the difference claimed.
    • Stopping the moment a variant looks ahead.
    • Changing two elements and attributing the result to the one you find more interesting.
    • Judging body-copy tests on open rate, which the body copy cannot influence.
    • Comparing two campaigns sent on different days and calling it a test.
    • Ignoring the unsubscribe column entirely.
    • Testing button colors while never testing the offer.

    Testing sits on top of fundamentals — if your list is stale or your authentication is broken, no test result will be stable enough to learn from. The email marketing overview covers the groundwork that makes testing meaningful in the first place.

    Frequently asked questions

    How large does my list need to be to A/B test emails?

    You can test with a few thousand contacts if you are testing bold, structural changes that would produce a large effect. Detecting differences of two points or less generally requires tens of thousands of recipients per arm.

    How long should I run an email A/B test?

    At least 24 hours, and 48 is safer. Engagement arrives in a long tail, and variant rankings often reverse between the first few hours and the second day.

    Can I test more than one thing at a time?

    Not in a standard A/B test — two changes leave you unable to explain the result. Multivariate testing can handle several variables at once, but it needs far more volume than most senders have.

    Should I judge subject line tests on open rate?

    Use open rate as a directional signal only. Image pre-fetching inflates opens unevenly, so confirm the winner with click rate or conversions before adopting it as a rule.

    What's a good win rate for email tests?

    Most well-designed tests produce no detectable difference, and that is normal. If nearly every test you run shows a winner, your samples are probably too small and you are reading noise as signal.

    Want built-in A/B testing that randomizes, splits, and sends the winner automatically? Launch on IGSendMail and run your next test properly from the first send.

    Enjoyed this article?

    Get email marketing tips delivered to your inbox every week.