The campaign went out at 9 a.m. on a Tuesday. Version A landed at a 31% open rate. Version B held steady at 24%. The marketing lead marked Version A the winner, updated the swipe file, and moved on. Six weeks later, the same team was puzzled: their list was engaged but their revenue wasn’t moving. Nobody connected those two facts, because the split test said everything was fine. It wasn’t fine. The open rate had been measuring the subject line’s ability to trigger curiosity, not its ability to attract the kind of reader who would ever buy. Email subject line testing is the most practiced ritual in email marketing and one of the most reliably misread signals in a small team’s toolkit.
Here’s the problem in one sentence: open rate measures a subject line’s appeal to everyone on your list, including the people who will never convert. A compelling subject line pulls in browsers, curious scrollers, and accidental subscribers just as effectively as it pulls in your actual buyers. When you optimize for opens, you’re optimizing for reach within a list, not for revenue. The two overlap less than most email marketers assume.
What the open rate actually measures
Open rate has one job: it tells you whether your subject line was interesting enough to trigger a pixel fire on a phone screen at 8:47 a.m. That’s it. It says nothing about fit, intent, or downstream behavior. It doesn’t distinguish between someone who opened, read every word, and bought, and someone who opened, glanced at the first sentence, and deleted it.
Before Apple’s Mail Privacy Protection rolled out broadly in late 2021, this was a flawed but functional proxy. MPP started pre-loading email pixels for Apple Mail users regardless of whether a human actually opened the message. Litmus research has tracked Apple Mail’s market share sitting above 55% of all email opens, which means more than half your open count may now include a machine pre-fetch, not a human eyeball. The metric was imperfect before MPP. Post-MPP, it’s an estimate wearing a confidence costume.
None of this means open rates are useless. They still catch deliverability problems. A sudden drop almost always points to spam-folder placement, not a bad subject line. But using them as the primary optimization target in a split test is a different matter entirely.
The variant that wins on opens often loses on revenue
Consider what each type of subject line actually does to audience composition. A curiosity gap subject line (“You’re probably doing this wrong…”) pulls opens from everyone who finds the ambiguity compelling. That’s a broad, intent-agnostic group. A specific, benefit-forward subject line (“How to reduce churn using your existing HubSpot data”) pulls opens from a narrower group: people who have HubSpot, who care about churn, and who are actively looking for solutions. The first line will almost certainly win a standard A/B open-rate test. The second line will almost certainly drive more revenue.
This isn’t hypothetical. The pattern is well-documented in conversion research. MarketingExperiments found in their email subject line research that specificity in subject lines consistently outperforms curiosity-based approaches on downstream conversion metrics. Curiosity lines optimize the top of the funnel within the email itself. Specific lines pre-qualify readers before they even open.
The mechanism matters: a specific subject line tells potential readers exactly what they’re opting into. Uninterested people opt out. That’s not a failure. That’s filtering. Your open-rate-optimized A/B test reads it as failure because the denominator went up and the numerator didn’t keep pace.
Email subject line testing: what to measure instead
Run your split tests. They’re still worth doing. Just change the dependent variable.
The metric hierarchy for a commercial email list looks like this, ordered by how directly it connects to business outcomes:
- Revenue per email sent, the most honest number. Divide total attributed revenue by total emails delivered. Noisy on small lists, but the signal you actually want.
- Click-to-open rate (CTOR), clicks divided by opens. This filters out the noise by looking only at readers who opened, then asking whether the email itself was relevant enough to drive action. A high open rate with a low CTOR is a curiosity subject line attracting the wrong crowd.
- Click rate on the list, total clicks divided by delivered, not just opens. Less susceptible to MPP inflation than open rate, and a more direct measure of engagement.
- Unsubscribe rate by variant, often overlooked, genuinely useful. A subject line that generates more unsubscribes is telling you it attracted readers who weren’t a fit, then disappointed them. That’s worth knowing.
Open rate still belongs on your dashboard. It just shouldn’t be the tiebreaker in a subject line test.
The sample size problem nobody talks about
Here’s a second failure mode layered on top of the first. Most small-team A/B tests on email don’t reach statistical significance before the sender calls a winner.
A typical scenario: a list of 4,000 subscribers, 20% sent to each variant (800 per arm), results checked after 24 hours. Version A: 248 opens. Version B: 210 opens. Winner declared. The problem is that a difference of 38 opens on a sample of 800, roughly a 5-point difference in open rate, requires a much larger sample to confidently attribute to the subject line rather than to random variation in who happened to check email that morning.
Tools like Klaviyo, Mailchimp, and ActiveCampaign all offer built-in A/B testing, and they all show you the winning percentage or confidence interval. But they default to showing you that number immediately, and it’s tempting to act on a 60% confidence reading as if it were 95%. It’s not. At 60% confidence, you’d be wrong about a third of the time if you ran the same test repeatedly.
The practical fix: before running a test, decide your minimum detectable effect. That’s the smallest difference you’d actually change your strategy over. Check whether your list size supports it. For most teams with lists under 10,000, a meaningful subject line test takes 3 to 5 separate sends to the same variant cohorts before the signal stabilizes. That’s not how most people run tests, but it’s the only way to trust the results.
Where this connects to your onboarding sequence
The open-rate trap compounds in email sequences. If you’re testing subject lines in an onboarding flow and optimizing for opens on Day 1, you may be pulling in curiosity-openers who disengage by Day 3. That’s exactly the failure pattern that kills onboarding flows. A subject line that pre-qualifies intent on Day 1 will sometimes look worse on opens and dramatically better on activation rate. Your onboarding email sequence’s success metric should be activation, not open rate. That means your subject line test should point toward the same destination.
The same logic extends to cold outreach. The cold email literature is unambiguous: reply rate, not open rate, is what actually predicts revenue from cold email. A subject line that gets 60% opens and 1% replies has beaten a subject line with 40% opens and 4% replies on the wrong metric.
The pre-qualifier test: a different way to write subject lines
Here’s a reframe worth naming. Instead of asking “which subject line gets more opens?”, ask “which subject line most accurately describes what’s inside, to the specific person who would benefit from it?”
Call this the pre-qualifier test. Before writing your subject line variants, run through three quick steps:
- Write one sentence describing your ideal reader for this email. What do they do? What problem do they have? What do they already believe?
- Ask whether your subject line would attract that person specifically, or attract anyone who finds the concept vaguely interesting.
- Score each variant. If the subject line could apply to three different buyer personas, it’s a curiosity line. Rewrite it until it only applies to one.
A curiosity-gap line like “The mistake 80% of marketers make” attracts the curious. A pre-qualifier line like “Why your HubSpot pipeline hides your actual churn risk” attracts HubSpot users thinking about churn. The second line fails the broad-audience open-rate test and passes the revenue test. Run the pre-qualifier test before you set up the A/B, not after you look at the results.
This reframe also changes how you write the email body. Once you’ve pre-qualified the reader in the subject line, you can write to them specifically. Not to a general audience. That specificity is what drives CTOR, what drives clicks, and what drives sales. The subject line and the body are one continuous argument, not two separate jobs.
One more thing the test won’t tell you
A split test compares two options against each other. It won’t tell you whether either option is good. A 31% vs. 24% open rate test reveals a relative winner, but both variants could be pulling unqualified openers at scale. The test result is always conditional on the quality of the hypotheses you started with.
This is where most testing programs plateau. The team runs tests, accumulates a swipe file of “winning” subject lines, and applies those patterns to future campaigns. But if the winning patterns were selected for open rate, the swipe file is a collection of curiosity-gap templates optimized for the wrong outcome. The patterns compound over time, and the gap between engaged opens and actual revenue quietly widens.
The fix is upstream: decide what winning means before the test runs. Revenue per send, CTOR, or reply rate. Any of these beats open rate as the north star. Once you’ve picked the right metric, the test results mean something. Until then, you’re declaring winners in a race where nobody checked which direction the finish line was.
