Cold Email Infrastructure

    Spam Filter Test: What Actually Triggers It

    Two different products are sold as a spam filter test. One scores a message, one samples a seed panel, and neither can see the layer that usually decides B2B placement.

    August 13, 20268 min read
    Share:
    The short answer

    A spam filter test is either a content and configuration check that scores one message against a rule engine, or a placement test that sends to a vendor-controlled seed panel and reports where copies landed. Both are useful for catching configuration errors. Neither can read your sending reputation or your recipients' own gateways.

    Key takeaways

    • Content tests score a message against a rule engine that is not your recipient's. Placement tests sample a panel that is not your audience. The two answer different questions.
    • A clean score alongside a spam landing is not a contradiction. Mailbox providers publish no equivalent score, so a rule engine's verdict is not a placement result.
    • Seed panels model the mailbox layer and not the organisational gateway sitting above it, which is where a large share of B2B mail is actually filtered.
    • The readouts about your own domain, such as spam rate and reputation in Postmaster Tools, beat any external score because they describe your sending rather than one sample message.

    Reviewed and updated August 13, 2026

    A campaign scored 9.8 out of 10 on a public spam test on Monday and landed in spam at two of the client's three largest accounts on Tuesday. Nothing was wrong with the test. It measured what it measures, reported it accurately, and had no way of knowing the thing that actually decided placement.

    Spam filter tests are worth running. They are also the most commonly over-read artifact in deliverability, because a single number invites the belief that it is a forecast. Knowing which of two very different tests you have run, and what each one can see, is most of the value.

    Two tools with one name

    Search for a spam filter test and you get two categories of product presented identically. They answer different questions and they fail in different ways.

    A content and configuration test takes one message, scores it against a rule engine, checks your authentication records, and often checks whether your sending host appears on public blocklists. You typically send one email to a generated address and read a report. The output is a score.

    A placement test sends to a panel of addresses the vendor controls across several providers, then reports where each copy landed: inbox, spam, or a secondary tab. The output is a distribution across providers.

    The first tells you about the message. The second tells you about a message's fate at a sample of mailboxes. Neither tells you where your actual campaign will land for your actual recipients, and the reasons differ.

    Content and configuration testOne message, one score
    • Scores the body against a rule engine's text rules
    • Checks SPF, DKIM and DMARC records resolve
    • Checks the sending host against public blocklists
    • Flags formatting: image-to-text ratio, link count, missing plain text
    • Cheap, instant, repeatable
    Placement testA panel of seed addresses
    • Reports inbox, spam or tab per provider
    • Covers the consumer providers well
    • Shows differences between providers you cannot otherwise see
    • Says nothing about your recipients' own gateways
    • The panel is not your audience
    The two products sold as a spam filter test, and the question each answers.

    What a score is

    A spam score is one rule engine's estimate that a message is unwanted, measured against a threshold that engine's operator chose. The major mailbox providers publish no equivalent number. That is the whole reason a clean score and a spam landing are not a contradiction.

    The engine behind most content tests is Apache SpamAssassin or something modelled on it. Its project page describes a "robust scoring framework" running "heuristic and statistical analysis tests on email headers and body text including text analysis, Bayesian filtering, DNS blocklists, and collaborative filtering databases" (spamassassin.apache.org). A public test can run the text rules and the blocklist lookups against your message honestly. It cannot run the Bayesian half in any way that resembles your recipient's installation, because that half is trained on mail that particular system has already seen.

    So the score is a real measurement of a real thing. The thing is narrower than the name suggests.

    What neither test can see

    Three inputs decide placement for B2B cold email, and no external test has access to any of them.

    Your reputation with that provider. Sender reputation is the receiving provider's accumulated judgment of your domain and your sending IP, built from what happened the last several thousand times mail from you arrived. A test message sent today carries that history with it, but the report cannot show you the history, only the outcome for that one message.

    Your recipients' own filtering. A large share of B2B mail never reaches Gmail or Microsoft's consumer stack at all. It reaches a corporate gateway, and the receiving organisation configures that gateway. This is the layer our spam filter definition separates out explicitly, and it is the layer that silently breaks placement tests: a seed panel does not sit behind your prospect's security appliance, so it cannot report what that appliance did.

    Engagement history. Whether your previous mail to similar recipients was opened, replied to, deleted unread or marked as spam feeds the next decision. A seed address has no behaviour, so it generates no engagement signal in either direction.

    The measurement that is actually available to you

    Google publishes real numbers about your own sending, and they are more useful than any external score because they describe your domain rather than a sample message. Google's sender guidelines require senders to keep spam rates reported in Postmaster Tools below 0.3%, and its Postmaster guidance advises keeping the reported rate below 0.10% while avoiding ever reaching 0.30% or higher (Google Workspace Admin Help).

    Those two numbers appear on the same page and are not in conflict. One is the requirement. The other is the operating target that leaves room for a bad week. Reading only the requirement is how a sender ends up parked one complaint away from a problem.

    The same page requires SPF or DKIM on sending domains, valid forward and reverse DNS records, TLS in transport, and message formatting per RFC 5322. Every one of those is verifiable from your own side before you send anything, which makes them a better pre-launch checklist than any score. Our Postmaster Tools guide covers reading the reputation and spam-rate panels, and the authentication guide covers the records themselves.

    1. Step 1Records

      Confirm SPF, DKIM and DMARC resolve and align for the sending domain, and that forward and reverse DNS exist for the host.

    2. Step 2Blocklists

      Check the domain and the sending IP against the public lists before the first send, not after a drop.

    3. Step 3Content test

      Run one message through a scoring tool. Treat anything it flags as an editing prompt rather than a verdict.

    4. Step 4Small live send

      Send to a modest slice of the real list from the real infrastructure and read replies, bounces and complaints.

    5. Step 5Provider readouts

      Watch domain reputation and spam rate for your own domain rather than a sample message's score.

    A pre-launch order that puts the cheap, decisive checks first.

    Why the small live send outranks the test

    The step most teams skip is the one that carries the most information. A test message tells you about a message. A modest send to real addresses on your real list tells you about the combination of your infrastructure, your list and your copy, which is the only combination that will ever ship.

    It also surfaces failures a test structurally cannot produce. A test address does not bounce, so it cannot tell you your list has decayed. A test address never sat on a spam trap, so it cannot warn you that your list was built in a way that will get you blocklisted. A test address does not have a colleague who reports you.

    Keep the slice small, keep it representative rather than cherry-picked, and read the outcome over a couple of days rather than an hour. Then scale. Our deliverability audit sets out the fuller version of this for an established sender, and the deliverability guide covers the repair path when the readouts come back bad.

    The seed-panel problem, stated plainly

    Placement tests deserve their own caveat because the number they produce looks so much like the number you want. A panel reports that 84% of copies reached the inbox. It is tempting to read that as "84% of my campaign will reach the inbox", and the two figures are not the same measurement.

    Our own definition of inbox placement makes the underlying point: nothing in the mail protocol reports where a delivered message landed, so every placement figure is an estimate produced by sampling addresses somebody controls. The estimate is only as representative as the panel.

    For consumer-heavy sending, panels are reasonably representative, because a large share of recipients really are at the handful of providers a panel covers. For B2B outbound the gap is wider. Your list is weighted toward corporate domains running Microsoft 365 or Google Workspace behind an organisation's own policies, and often behind a third-party gateway in front of that. A panel of vendor-owned mailboxes at the major providers models the mailbox layer and not the organisational layer sitting above it.

    The engagement gap is the other half of it, and our deliverability platforms comparison covers that ground in detail alongside the vendor categories, so the short version is enough here: a seed address brings no behavioural history, which usually makes a seed result the more pessimistic of the two readings. That page is also where to go if the question you actually have is which category of tool to buy, since placement testing, warm-up and monitoring are three different products routinely sold under one phrase.

    None of this makes the panel useless. A split result across providers is real information, and a panel is the only practical way to see it. Read it as a comparison between providers rather than as a percentage you can apply to your own send.

    The case worth planning for is the two tests disagreeing, because it is common and it is informative rather than confusing. A clean content score with a poor placement result points away from the message and toward reputation, list quality or a receiving gateway. A poor content score with good placement usually means the rule engine flagged a register the providers currently tolerate on your domain, which is a warning about headroom rather than a present failure. The combination that should stop a launch is a poor result on both, since that is the one case where the two instruments agree and neither is being asked to see past its own range.

    Reading a bad result honestly

    When a test comes back poor, the useful discipline is to separate findings you can act on from findings that are noise.

    Authentication failures, missing DNS records and blocklist hits are unambiguous. Fix them; they are facts about your configuration, and they are the same fact for every recipient.

    Content flags are a register check. A rule engine adding points for promotional phrasing has told you the copy reads promotional, which is worth knowing and is not the same as a placement prediction. Rewriting to remove flagged tokens while keeping the register changes the score without changing the outcome.

    A poor placement result at one provider and a clean result at another is normal and informative. Providers weigh signals differently, and a split result usually points at reputation with one of them rather than at the message.

    A clean score alongside a real placement problem means the problem is in a layer the test does not reach, which is almost always reputation, list quality, or the receiving organisation's own gateway.

    The short version

    Run the tests, and read them as instruments with a known range. Content tests measure a message against a rule engine that is not your recipient's. Placement tests measure a panel that is not your audience. Both are worth the few minutes they cost, particularly for catching configuration errors before launch. Neither replaces the readouts about your own domain, or the small live send that exercises every layer at once. If you want that diagnosed against real sending rather than a sample message, our free campaign build is where we do it.

    Questions

    Frequently asked questions.

    Frequently asked questions
    Why did my email score 10 out of 10 and still land in spam?
    Because the score measured the message and the filter judged the sender. A content test runs text rules, checks your authentication records and looks up blocklists. It cannot see your domain's reputation with that provider, your recipients' engagement history, or a corporate gateway sitting in front of the mailbox. All three outrank wording.
    Are seed list placement tests accurate for B2B outbound?
    Less than for consumer sending. Your list is weighted toward corporate domains running their own policies, often behind a third-party gateway, and a panel of vendor mailboxes does not sit behind those. Seed addresses also accumulate no genuine engagement. Read a panel result as a comparison between providers rather than a percentage for your own send.
    What should I check before launching a campaign?
    Confirm SPF, DKIM and DMARC resolve and align, that forward and reverse DNS exist for the sending host, and that the domain and IP are absent from public blocklists. Then run one message through a scoring tool as an editing prompt. Then send to a small representative slice of the real list and read bounces, replies and complaints.
    Which spam test result should I actually act on?
    Authentication failures, missing DNS records and blocklist hits are unambiguous facts about your configuration and are the same for every recipient, so fix those first. Content flags are a register check rather than a prediction. A split result across providers usually points at reputation with one of them rather than at the message.
    Email DeliverabilityCold EmailSpam FiltersInbox PlacementEmail Infrastructure
    Byline

    About the author.

    Tim Carden

    Tim Carden is CMO / CTO at RevenueFlow, which builds and operates outbound revenue engines for B2B companies. Studied at McGill University.

    Tim Carden · CMO / CTO

    Connect on LinkedIn →
    Your next move

    Ready to scale your outreach?

    We build GTM engines that book real meetings. See the receipts.

    Further reading

    Related articles.

    Cold Email Infrastructure

    Spam Filter Barracuda: How to Diagnose It Before It Costs a Domain

    Two Barracudas sit between a cold email and a B2B inbox, and only one has a lookup. Written for the sender being blocked rather than the administrator doing it.

    10 min readRead →
    Cold Email Infrastructure

    How to Bypass a Spam Filter: The Only Method That Works, and Who Holds It

    One reliable bypass exists and the recipient's administrator holds it. What the admin controls actually do, and why sender-side bypass tactics make placement worse.

    7 min readRead →
    Cold Email Infrastructure

    Office 365 Spam Filter: What Actually Triggers It

    The same message lands in the inbox at one Microsoft tenant and in quarantine at the next. A setting the recipient chose decides which, and the headers say so.

    7 min readRead →
    Cold Email Infrastructure

    Bounce Back Email for B2B Teams: How to Diagnose It Before It Costs a Domain

    A bounce back email carries the receiving server's verbatim reason for refusing you. Here is how to read the codes, and what the shape of a batch tells you.

    9 min readRead →
    Cold Email Infrastructure

    Spam Word: Diagnosing Placement Without Guesswork

    Spam word lists describe the text rules inside a scoring engine, and a scoring engine's verdict is not a placement result. Where vocabulary actually earns weight.

    7 min readRead →
    Cold Email Infrastructure

    Google Spam Filter for B2B Teams: Diagnosing Placement Without Guesswork

    Google publishes what it wants from senders, and the list is short and checkable. What binds a B2B sender, which spam rate to watch, and what to do when placement drops.

    7 min readRead →