You build the campaign in Klaviyo, run it through a spam checker as a last step before scheduling, and it comes back with a 9 out of 10. Authentication passes. No obvious red flags in the copy. You hit send.
Two hours later you’re looking at open rates that are half of what this segment usually gets, and a quick seed check confirms it: a meaningful share landed in Gmail’s spam folder anyway.
The score wasn’t wrong, exactly. It was just answering a narrower question than the one you were actually asking. InboxEagle’s inbox placement monitoring infrastructure tracked 528,475 real ecommerce marketing messages between June 11 and September 12, 2026, from brands actually sending through Klaviyo and other ESPs, not test sends. Each message was scored, risk level and all, from its own placement, content, and authentication, and separately recorded for exactly where it landed: inbox, promotions, or spam. We looked at how those scores lined up against that real outcome. Some of it confirms what you’d expect. Some of it doesn’t.
528,475 Monitored Ecommerce Messages, Jun 11–Sep 12, 2026
Testing Your Email’s Spam Score: What the Risk Levels Actually Predict
When you check your email’s spam score with an email content spam checker, it typically buckets the message into a risk tier: low, medium, high, critical. The assumption baked into that label is that the tiers rank in order, and that a “high” risk message is meaningfully more likely to hit spam than a “medium” or “low” one.
Checked against actual Gmail placement, that assumption only holds at one end of the scale.
Risk Tier vs. Gmail Placement (454,989 messages)
A message flagged “critical” landed in spam 36.5% of the time, roughly 1.7x the rate of a “high” risk message at 21.8%. That part tracks. But “high” risk mail had a lower spam rate than both “medium” (25.9%) and “low” (26.2%). If the tiers ranked cleanly, that shouldn’t happen. The honest read is that the score cleanly separates one thing, an obvious critical-severity problem, from everything else. The gradations between high, medium, and low don’t reliably predict which one is safer.
There’s a second wrinkle worth flagging rather than smoothing over: critical-risk mail didn’t perform worse across the board. It had the highest inbox rate of any tier too, at 12.4%, and by far the lowest promotions rate, at 51.1% versus 64.5–66.4% everywhere else. Critical-risk mail is polarized. It’s far less likely to sit in Gmail’s Promotions tab and far more likely to land at one extreme or the other. That’s a different story than “critical mail is universally worse,” and it matters for how you read your own results. How spam filters actually score a message covers the categories a checker like this one is built from.
The risk-tier mix itself also isn’t uniform by provider, even though only Gmail’s cut breaks placement down by folder in this dataset. This next comparison uses total volume per risk tier, which all three providers report, rather than folder outcome, which only Gmail does.
Risk-Tier Mix by Provider (528,475 messages)
| Provider | Messages | Critical | High | Medium | Low |
|---|---|---|---|---|---|
| AOL | 32,761 | 0.1% | 25.6% | 38.3% | 36.1% |
| Gmail | 454,989 | 0.1% | 21.7% | 38.4% | 39.8% |
| Yahoo | 40,725 | 0.1% | 19.3% | 36.0% | 44.7% |
AOL-bound mail carried the highest “high” risk share of the three providers, at 25.6%, against 19.3% at Yahoo. Yahoo’s volume skewed the other way, with 44.7% landing in the “low” tier compared with 36.1% at AOL. The critical-risk share sits at roughly 0.06–0.07% everywhere, close enough to call it the same number across all three. It’s also the one figure in this section that’s as reliable for AOL and Yahoo as it is for Gmail, since it’s a share of total volume rather than a folder-placement outcome those two providers don’t report in this cut.
The One Thing That Actually Creates a “Critical” Flag
If the risk tiers only cleanly separate at the top, the next question is what actually pushes a message into that top tier. Across the same dataset, the answer is almost entirely one thing: authentication.
Messages that were not fully authenticated made up 13,102 of the 528,475 messages monitored, just 2.5% of volume. That same 2.5% accounted for every single message flagged critical risk. There is no “critical” row for fully authenticated mail in this dataset at all, because a message with SPF, DKIM, and DMARC all passing never scored high enough to reach it.
| Authentication | Risk Level | Messages | Avg. Score |
|---|---|---|---|
| Not fully authenticated | Critical | 331 | 54.41 |
| Not fully authenticated | High | 4,971 | 36.16 |
| Fully authenticated | High | 109,882 | 39.37 |
| Not fully authenticated | Medium | 7,800 | 25.04 |
| Fully authenticated | Medium | 194,263 | 24.19 |
| Fully authenticated | Low | 211,228 | 16.38 |
Read the medium row and the pattern holds the way you’d expect: broken authentication nudges the average score up, from 24.19 to 25.04. But the high-risk row does the opposite. Not-fully-authenticated messages in the high tier averaged 36.16, lower than the 39.37 average for fully authenticated high-risk messages. That’s worth stating honestly rather than papering over: authentication failures don’t uniformly raise the numeric score within every tier, but they are the sole reason the critical tier exists at all. Below a certain severity, a broken record just clears the bar for “high.” Past it, it’s the only thing that clears the bar for “critical.” Our DMARC failure study across 2.2 million emails covers the placement penalty from failed authentication in more depth, and SPF, DKIM, and DMARC explained covers what each record actually checks.
The Score Barely Differs Between Inbox and Spam Outcomes
This is the finding that should change how you use a spam score, not just what you check with it. Placement is one of the things InboxEagle’s engine can see when it scores a message, alongside content and authentication. Even so, the score barely tracks the message’s own outcome.
| Actual Placement | Messages | Avg. Content Penalty | Avg. Final Score |
|---|---|---|---|
| Inbox | 118,387 | 35.11 | 24.93 |
| Promotions | 295,794 | 33.66 | 24.41 |
| Spam | 114,294 | 30.71 | 23.69 |
The average final score for messages that landed in spam was 23.69, lower than the 24.93 average for messages that landed in the inbox. On a scale where a critical-risk message averages 54.41, a 1.24-point gap between the spam bucket and the inbox bucket is close to flat. A message that scored slightly worse was not meaningfully more likely to have landed in spam than one that scored slightly better.
One limitation to be direct about: the “inbox” bucket above also absorbs all of the AOL and Yahoo volume from this cut, since those two providers’ placement outside inbox isn’t tracked in this dataset. That means the inbox bucket isn’t a clean Gmail-only comparison the way the risk-tier breakdown earlier is. Even accounting for that, a compressed 1.24-point spread between a 118,387-message bucket and a 114,294-message bucket is a small gap either way you slice it.
This isn’t proof the score is useless. It’s evidence that authentication and content signals dominate the score far more than the campaign’s own placement outcome does. A composite score like this can see your authentication setup and your content. What it can’t fully capture is the reputation and engagement history of the domain sending to that specific list, which is exactly what decides placement most, and exactly the part seed list testing vs. real-world inbox monitoring exists to explain. This is population-level scoring across thousands of brands, not a forecast for your next individual send.
Which Words Your Spam Score Checker Actually Penalizes
Content checkers score individual words and phrases against a rule set, and it’s worth knowing what those rules actually reward and punish, separate from whether they predict your real outcome.
Think of these as spam-test indicators, not verdicts: signals a checker matches on, not proof of what will happen to your campaign.
| Phrase | Occurrences | Spam Placement Rate |
|---|---|---|
| “#1” | 8,767 | 30.32% |
| “special offer” | 10,517 | 29.76% |
| “the best” | 17,019 | 23.50% |
| “cure” | 34,758 | 22.96% |
| “expire” | 11,964 | 22.14% |
| “cannot be combined with any other offer” | 10,277 | 9.89% |
| “credit card” | 20,457 | 6.74% |
| “bank” | 12,979 | 9.72% |
| “credit” | 34,443 | 10.97% |
| “cash” | 45,778 | 11.14% |
“#1” and “special offer” carried the highest phrase-level spam rates in the entire dataset, at 30.32% and 29.76%. But “credit card” and “bank”, two of the phrases most spam-word checklists flag first, sit at the bottom, at 6.74% and 9.72%. Neither of those low numbers means those words are “safe” in some universal sense. It means they mostly show up in legitimate transactional and financial mail in this dataset, and a rule-based checker scores the phrase, not the context it appears in. This tracks with what our 774,828-email trigger word study found at a broader level: word choice correlates with placement far less than who the message is actually going to. Use the spam word checker as a proofreading pass on a specific subject line, not as a placement forecast.
Subject Line Length: What It Actually Does to Placement
The last cut looks at subject line length against actual placement directly, across a broader pull of 598,544 messages from the same window.
Subject Line Length vs. Placement
Subject lines in the 80–100 character band had both the highest spam rate in the dataset, at 27.24%, and the highest inbox rate, at 33.21%. Length doesn’t reliably improve or hurt placement here. It mostly pulls mail out of Gmail’s Promotions tab: the promotions share drops from 59.45% for subject lines under 20 characters to 39.55% in the 80–100 character band, and that volume splits toward both inbox and spam instead. Average exclamation mark usage also climbed alongside length in this cut, from 0.15 at the shortest band to 0.40 at the longest, so length and punctuation are confounded here and this data can’t isolate which one is doing the work. Treat this as a reason to check both outcomes on a longer subject line, not a reason to avoid one. Subject line personalization and inbox placement covers a related variable from the same angle, and the subject line analyzer is a faster way to test a specific line than waiting for send data.
So What Do You Actually Do Before You Hit Send in Klaviyo
- Treat a “critical” flag as a hard stop, not a spectrum. It’s the only tier in this dataset that cleanly separates from the rest, at a 36.5% Gmail spam rate against 21.8–26.2% for everything else. Don’t spend equal effort chasing a “high” score down to “medium.”
- Fix authentication before you touch a word of copy. It’s 2.5% of volume and 100% of critical flags. Verify SPF, DKIM, and DMARC on the exact sending domain in Klaviyo’s sending domain settings before reviewing anything else.
- Don’t treat a passing score as a placement guarantee for your specific send. Messages that landed in spam scored 23.69 on average; messages that landed in the inbox scored 24.93. That 1.24-point gap held across thousands of brands’ worth of real campaigns. A good score reflects clean authentication and content. It isn’t a forecast for where any one campaign lands.
- Use word-level flags as a proofreading pass, not gospel. “Credit card” and “bank” scored lower than “the best” and “#1” in this dataset. A checker flags rule matches, not risk.
- Test subject line length deliberately, and watch both ends. A longer subject line here raised inbox and spam rates together. Check where it lands, not just whether it scored well.
- Confirm real placement with a seed list before you send to your full list. A one-off check on your own campaign is a narrower, static read than InboxEagle’s ongoing monitoring. A seed list, the closest thing on the market to a real email inbox tester, shows you what actually happened at Gmail, Yahoo, and Outlook within minutes, on this specific campaign.
The Bottom Line
Across 528,475 monitored ecommerce campaigns, a critical-risk flag meant something real: a 36.5% Gmail spam rate against 21.8–26.2% for every other tier, traced almost entirely to broken authentication, itself only 2.5% of tracked volume. But even though InboxEagle’s scoring engine can see where a message actually landed, the score barely tracked that outcome: messages that landed in spam scored, on average, no worse than messages that landed in the inbox. A composite spam-risk score built from authentication and content signals is a real, useful health check for a sending domain. It is not a placement guarantee for any single campaign, because population-level scoring can’t capture the reputation and engagement history that decides where one specific send actually goes.
528,475 spam tests analyzed
A passing spam score isn't a placement guarantee. Check both.
Run your Klaviyo sending domain through InboxEagle's free DMARC checker to confirm SPF, DKIM, and DMARC all pass, and see live inbox, promotions, and spam placement for that domain in the same report. Then run a seed list test to check a specific campaign before it goes to your full list.
Methodology
Data period: June 11–September 12, 2026, the same window as our companion studies on Gmail spam placement and DMARC enforcement.
Data source: InboxEagle’s inbox placement monitoring infrastructure, which tracks real ecommerce marketing campaigns from thousands of monitored brands as they’re actually sent (our companion Gmail study covered mail from 17,524 brands in this same window). There is no separate pre-send test step in this dataset. Each message was scored for spam risk from its own placement outcome, content, and authentication together, and that same message’s actual inbox, promotions, or spam placement was recorded alongside it. The score and the placement are two measurements of the same real, already-sent campaign, not a prediction checked against a later result.
Sample sizes: 528,475 messages in the risk-level, authentication, and content-penalty cuts (consistent across all three, cross-checked by summing each breakdown). Of those, 454,989 were sent through Gmail, 40,725 through Yahoo, and 32,761 through AOL. The subject line length cut draws from a separate pull of 598,544 messages from the same window and doesn’t reconcile to the 528,475 figure above; treat it as its own sample rather than a subset. The phrase-level spam word cut covers 51 tracked phrases with individual occurrence counts and isn’t a sample size in the same sense, since a single message can contain multiple tracked phrases or none.
Limitations, stated plainly: Because placement is itself one input to the score, the score-vs-placement comparisons in this post show how much authentication and content signals dominate the score relative to a campaign’s own outcome, not a blind prediction being checked against reality. AOL and Yahoo only report whether a message reached an inbox in this cut, not promotions vs. spam, consistent with our companion Gmail study, so every spam/promotions rate above uses the Gmail subset unless noted. This is a monitored panel skewed toward ecommerce brands using major ESPs, not a representative census of all email.
