Spam Filter Test: What Actually Triggers It
Two different products are sold as a spam filter test. One scores a message, one samples a seed panel, and neither can see the layer that usually decides B2B placement.
A spam filter test is either a content and configuration check that scores one message against a rule engine, or a placement test that sends to a vendor-controlled seed panel and reports where copies landed. Both are useful for catching configuration errors. Neither can read your sending reputation or your recipients' own gateways.
Key takeaways
- Content tests score a message against a rule engine that is not your recipient's. Placement tests sample a panel that is not your audience. The two answer different questions.
- A clean score alongside a spam landing is not a contradiction. Mailbox providers publish no equivalent score, so a rule engine's verdict is not a placement result.
- Seed panels model the mailbox layer and not the organisational gateway sitting above it, which is where a large share of B2B mail is actually filtered.
- The readouts about your own domain, such as spam rate and reputation in Postmaster Tools, beat any external score because they describe your sending rather than one sample message.
Reviewed and updated August 13, 2026
A campaign scored 9.8 out of 10 on a public spam test on Monday and landed in spam at two of the client's three largest accounts on Tuesday. Nothing was wrong with the test. It measured what it measures, reported it accurately, and had no way of knowing the thing that actually decided placement.
Spam filter tests are worth running. They are also the most commonly over-read artifact in deliverability, because a single number invites the belief that it is a forecast. Knowing which of two very different tests you have run, and what each one can see, is most of the value.
Two tools with one name
Search for a spam filter test and you get two categories of product presented identically. They answer different questions and they fail in different ways.
A content and configuration test takes one message, scores it against a rule engine, checks your authentication records, and often checks whether your sending host appears on public blocklists. You typically send one email to a generated address and read a report. The output is a score.
A placement test sends to a panel of addresses the vendor controls across several providers, then reports where each copy landed: inbox, spam, or a secondary tab. The output is a distribution across providers.
The first tells you about the message. The second tells you about a message's fate at a sample of mailboxes. Neither tells you where your actual campaign will land for your actual recipients, and the reasons differ.
- Scores the body against a rule engine's text rules
- Checks SPF, DKIM and DMARC records resolve
- Checks the sending host against public blocklists
- Flags formatting: image-to-text ratio, link count, missing plain text
- Cheap, instant, repeatable
- Reports inbox, spam or tab per provider
- Covers the consumer providers well
- Shows differences between providers you cannot otherwise see
- Says nothing about your recipients' own gateways
- The panel is not your audience
What a score is
A spam score is one rule engine's estimate that a message is unwanted, measured against a threshold that engine's operator chose. The major mailbox providers publish no equivalent number. That is the whole reason a clean score and a spam landing are not a contradiction.
The engine behind most content tests is Apache SpamAssassin or something modelled on it. Its project page describes a "robust scoring framework" running "heuristic and statistical analysis tests on email headers and body text including text analysis, Bayesian filtering, DNS blocklists, and collaborative filtering databases" (spamassassin.apache.org). A public test can run the text rules and the blocklist lookups against your message honestly. It cannot run the Bayesian half in any way that resembles your recipient's installation, because that half is trained on mail that particular system has already seen.
So the score is a real measurement of a real thing. The thing is narrower than the name suggests.
What neither test can see
Three inputs decide placement for B2B cold email, and no external test has access to any of them.
Your reputation with that provider. Sender reputation is the receiving provider's accumulated judgment of your domain and your sending IP, built from what happened the last several thousand times mail from you arrived. A test message sent today carries that history with it, but the report cannot show you the history, only the outcome for that one message.
Your recipients' own filtering. A large share of B2B mail never reaches Gmail or Microsoft's consumer stack at all. It reaches a corporate gateway, and the receiving organisation configures that gateway. This is the layer our spam filter definition separates out explicitly, and it is the layer that silently breaks placement tests: a seed panel does not sit behind your prospect's security appliance, so it cannot report what that appliance did.
Engagement history. Whether your previous mail to similar recipients was opened, replied to, deleted unread or marked as spam feeds the next decision. A seed address has no behaviour, so it generates no engagement signal in either direction.
The measurement that is actually available to you
Google publishes real numbers about your own sending, and they are more useful than any external score because they describe your domain rather than a sample message. Google's sender guidelines require senders to keep spam rates reported in Postmaster Tools below 0.3%, and its Postmaster guidance advises keeping the reported rate below 0.10% while avoiding ever reaching 0.30% or higher (Google Workspace Admin Help).
Those two numbers appear on the same page and are not in conflict. One is the requirement. The other is the operating target that leaves room for a bad week. Reading only the requirement is how a sender ends up parked one complaint away from a problem.
The same page requires SPF or DKIM on sending domains, valid forward and reverse DNS records, TLS in transport, and message formatting per RFC 5322. Every one of those is verifiable from your own side before you send anything, which makes them a better pre-launch checklist than any score. Our Postmaster Tools guide covers reading the reputation and spam-rate panels, and the authentication guide covers the records themselves.
- Step 1Records
Confirm SPF, DKIM and DMARC resolve and align for the sending domain, and that forward and reverse DNS exist for the host.
- Step 2Blocklists
Check the domain and the sending IP against the public lists before the first send, not after a drop.
- Step 3Content test
Run one message through a scoring tool. Treat anything it flags as an editing prompt rather than a verdict.
- Step 4Small live send
Send to a modest slice of the real list from the real infrastructure and read replies, bounces and complaints.
- Step 5Provider readouts
Watch domain reputation and spam rate for your own domain rather than a sample message's score.
Why the small live send outranks the test
The step most teams skip is the one that carries the most information. A test message tells you about a message. A modest send to real addresses on your real list tells you about the combination of your infrastructure, your list and your copy, which is the only combination that will ever ship.
It also surfaces failures a test structurally cannot produce. A test address does not bounce, so it cannot tell you your list has decayed. A test address never sat on a spam trap, so it cannot warn you that your list was built in a way that will get you blocklisted. A test address does not have a colleague who reports you.
Keep the slice small, keep it representative rather than cherry-picked, and read the outcome over a couple of days rather than an hour. Then scale. Our deliverability audit sets out the fuller version of this for an established sender, and the deliverability guide covers the repair path when the readouts come back bad.
The seed-panel problem, stated plainly
Placement tests deserve their own caveat because the number they produce looks so much like the number you want. A panel reports that 84% of copies reached the inbox. It is tempting to read that as "84% of my campaign will reach the inbox", and the two figures are not the same measurement.
Our own definition of inbox placement makes the underlying point: nothing in the mail protocol reports where a delivered message landed, so every placement figure is an estimate produced by sampling addresses somebody controls. The estimate is only as representative as the panel.
For consumer-heavy sending, panels are reasonably representative, because a large share of recipients really are at the handful of providers a panel covers. For B2B outbound the gap is wider. Your list is weighted toward corporate domains running Microsoft 365 or Google Workspace behind an organisation's own policies, and often behind a third-party gateway in front of that. A panel of vendor-owned mailboxes at the major providers models the mailbox layer and not the organisational layer sitting above it.
The engagement gap is the other half of it, and our deliverability platforms comparison covers that ground in detail alongside the vendor categories, so the short version is enough here: a seed address brings no behavioural history, which usually makes a seed result the more pessimistic of the two readings. That page is also where to go if the question you actually have is which category of tool to buy, since placement testing, warm-up and monitoring are three different products routinely sold under one phrase.
None of this makes the panel useless. A split result across providers is real information, and a panel is the only practical way to see it. Read it as a comparison between providers rather than as a percentage you can apply to your own send.
The case worth planning for is the two tests disagreeing, because it is common and it is informative rather than confusing. A clean content score with a poor placement result points away from the message and toward reputation, list quality or a receiving gateway. A poor content score with good placement usually means the rule engine flagged a register the providers currently tolerate on your domain, which is a warning about headroom rather than a present failure. The combination that should stop a launch is a poor result on both, since that is the one case where the two instruments agree and neither is being asked to see past its own range.
Reading a bad result honestly
When a test comes back poor, the useful discipline is to separate findings you can act on from findings that are noise.
Authentication failures, missing DNS records and blocklist hits are unambiguous. Fix them; they are facts about your configuration, and they are the same fact for every recipient.
Content flags are a register check. A rule engine adding points for promotional phrasing has told you the copy reads promotional, which is worth knowing and is not the same as a placement prediction. Rewriting to remove flagged tokens while keeping the register changes the score without changing the outcome.
A poor placement result at one provider and a clean result at another is normal and informative. Providers weigh signals differently, and a split result usually points at reputation with one of them rather than at the message.
A clean score alongside a real placement problem means the problem is in a layer the test does not reach, which is almost always reputation, list quality, or the receiving organisation's own gateway.
The short version
Run the tests, and read them as instruments with a known range. Content tests measure a message against a rule engine that is not your recipient's. Placement tests measure a panel that is not your audience. Both are worth the few minutes they cost, particularly for catching configuration errors before launch. Neither replaces the readouts about your own domain, or the small live send that exercises every layer at once. If you want that diagnosed against real sending rather than a sample message, our free campaign build is where we do it.
Frequently asked questions.
Frequently asked questions- Why did my email score 10 out of 10 and still land in spam?
- Because the score measured the message and the filter judged the sender. A content test runs text rules, checks your authentication records and looks up blocklists. It cannot see your domain's reputation with that provider, your recipients' engagement history, or a corporate gateway sitting in front of the mailbox. All three outrank wording.
- Are seed list placement tests accurate for B2B outbound?
- Less than for consumer sending. Your list is weighted toward corporate domains running their own policies, often behind a third-party gateway, and a panel of vendor mailboxes does not sit behind those. Seed addresses also accumulate no genuine engagement. Read a panel result as a comparison between providers rather than a percentage for your own send.
- What should I check before launching a campaign?
- Confirm SPF, DKIM and DMARC resolve and align, that forward and reverse DNS exist for the sending host, and that the domain and IP are absent from public blocklists. Then run one message through a scoring tool as an editing prompt. Then send to a small representative slice of the real list and read bounces, replies and complaints.
- Which spam test result should I actually act on?
- Authentication failures, missing DNS records and blocklist hits are unambiguous facts about your configuration and are the same for every recipient, so fix those first. Content flags are a register check rather than a prediction. A split result across providers usually points at reputation with one of them rather than at the message.
About the author.
Tim Carden is CMO / CTO at RevenueFlow, which builds and operates outbound revenue engines for B2B companies. Studied at McGill University.
Tim Carden · CMO / CTO
Connect on LinkedIn →Explore more.
Ready to scale your outreach?
We build GTM engines that book real meetings. See the receipts.
Related articles.
Spam Filter Barracuda: How to Diagnose It Before It Costs a Domain
Two Barracudas sit between a cold email and a B2B inbox, and only one has a lookup. Written for the sender being blocked rather than the administrator doing it.
How to Bypass a Spam Filter: The Only Method That Works, and Who Holds It
One reliable bypass exists and the recipient's administrator holds it. What the admin controls actually do, and why sender-side bypass tactics make placement worse.
Office 365 Spam Filter: What Actually Triggers It
The same message lands in the inbox at one Microsoft tenant and in quarantine at the next. A setting the recipient chose decides which, and the headers say so.
Bounce Back Email for B2B Teams: How to Diagnose It Before It Costs a Domain
A bounce back email carries the receiving server's verbatim reason for refusing you. Here is how to read the codes, and what the shape of a batch tells you.
Spam Word: Diagnosing Placement Without Guesswork
Spam word lists describe the text rules inside a scoring engine, and a scoring engine's verdict is not a placement result. Where vocabulary actually earns weight.
Google Spam Filter for B2B Teams: Diagnosing Placement Without Guesswork
Google publishes what it wants from senders, and the list is short and checkable. What binds a B2B sender, which spam rate to watch, and what to do when placement drops.