Cold Email Infrastructure

    Graymail vs Spam: What the Difference Costs a Sender

    Graymail is a claim about a sending pattern. Spam is a claim about intent. They respond to different repairs, and confusing them sends teams to rewrite copy.

    Editorial illustration for Graymail vs Spam
    September 2, 2026Updated September 2, 20267 min read
    Share:
    The short answer

    Graymail is bulk mail from a legitimate sender that a filter scores on sending pattern and complaint history rather than on intent. Spam is a claim about intent. The two verdicts come from different mechanisms, so a bulk placement is repaired by changing who is on the list rather than by rewriting subject lines.

    Key takeaways

    • Microsoft's Defender documentation states that messages from a bulk sender are known as bulk mail or gray mail, and publishes a scale from 0 to 9 scored on complaint volume.
    • The published default bulk complaint level threshold is 7, with 6 under the Standard preset security policy and 5 under the Strict preset, so the same message lands differently by recipient organisation.
    • Google's sender guidelines tell all senders to keep spam rates in Postmaster Tools below 0.3%, which is the number a cold programme can actually move.
    • Cold outreach is bulk by construction, so the achievable goal is the low end of the bulk scale rather than an exit from it.

    Reviewed and updated September 2, 2026

    A sender checks the seed inbox, finds the message in Promotions rather than Primary, and starts rewriting subject lines. Nothing in the copy caused it. The filter did not decide the message was abusive. It decided the message was bulk, and bulk has its own bucket with its own rules.

    Graymail is the name for that bucket. Understanding the difference between graymail and spam matters to a cold sender for one reason: the two verdicts are produced by different mechanisms, respond to different repairs, and confusing them sends teams to rewrite copy when the actual signal was engagement.

    What each word means to the people who filter mail

    The definitions worth using are the ones the filtering vendors publish, because those are the definitions their products implement.

    Proofpoint's own threat reference page on graymail, fetched on 2 September 2026, defines it as "bulk email that does not fit the definition of spam because it is solicited, comes from a legitimate source, and has varying value to different recipients". Its examples are newsletters, announcements and targeted advertisements, and it notes that recipients previously opted in, knowingly or unknowingly, while the value of the mailing may have decreased over time.

    Microsoft uses the same concept under a different name and makes the mechanism explicit. Its Defender for Office 365 documentation on bulk complaint level, fetched the same day, states that "Messages from a bulk sender are known as bulk mail or gray mail", and describes a numeric score attached to inbound messages in an X-header. The published scale runs from 0, where "The message isn't from a bulk sender", through 1 to 3 for "a bulk sender that generates few complaints", 4 to 7 for a mixed number of complaints, and 8 to 9 for "a bulk sender that generates a high number of complaints".

    Spam is the other category, and its defining property is not volume. It is that the message is unsolicited, frequently deceptive, and sometimes malicious. A filter reaching a spam verdict is making a claim about intent. A filter reaching a graymail verdict is making a claim about a sending pattern and a complaint history, and it is explicitly declining to make the intent claim.

    The verdict is a scale, not a switch

    This is the part that changes how a cold sender should read a placement problem.

    Microsoft's documentation publishes the thresholds directly. The default bulk complaint level threshold used in anti-spam policies is 7. The Standard preset security policy uses 6, and the Strict preset security policy uses 5. Messages meeting or exceeding the configured threshold have the policy's bulk action taken on them.

    Three things follow from that, and none of them is about your copy.

    The threshold is set by the recipient's administrator, not by you. The same message, from the same domain, on the same day, can land in the inbox at one company and in junk at another because one of them selected a stricter preset. A seed test in your own mailbox measures your own configuration.

    The score is driven by complaint history rather than by content. A bulk sender generating few complaints scores low. The lever is who you send to and whether they mind, which is a list decision made before anything was written.

    And there is a middle band that behaves like neither inbox nor spam. A message scored as bulk is not blocked and is not accused of anything. It is filed. On the consumer side the same middle band is what the Promotions tab does, and the corpus covers what that means for placement diagnosis in the spam folder, which sets out the three destinations rather than the usual two.

    Graymail, or bulkA claim about a pattern
    • Solicited or at least not obviously unwanted, from a real sender
    • Scored on complaint history and sending pattern
    • Filed into a bucket rather than blocked
    • Threshold chosen by the recipient's administrator
    • Repaired by changing who is on the list
    SpamA claim about intent
    • Unsolicited, often deceptive, sometimes malicious
    • Scored on content, authentication, reputation and known signatures
    • Blocked, quarantined or filed as junk
    • Thresholds far less discretionary
    • Repaired by fixing authentication, list quality and complaint rate
    Two verdicts, two mechanisms, two different repairs. Reading one as the other is what sends teams to rewrite copy that was never the cause.

    Where cold outreach actually sits

    Section illustration: Where cold outreach actually sits

    Worth accepting rather than arguing with: B2B cold outreach lives in the graymail neighbourhood by construction. It is bulk by any reasonable definition, it comes from a real sender, and its value genuinely varies by recipient. That is the honest description, and the corpus makes the same point when reading a gateway's own material in the Mimecast spam filter.

    The one place cold outreach differs from the newsletter case Proofpoint describes is consent. A newsletter recipient opted in at some point and later stopped caring. A cold recipient never opted in at all. That gap is what makes complaint rate the binding constraint on the whole motion, because complaint rate is what a bulk score is built from and it is the one number a cold programme can move.

    Google publishes its side of this as a hard requirement rather than as guidance. Its sender guidelines page, fetched on 2 September 2026, tells all senders to "Keep spam rates reported in Postmaster Tools below 0.3%", and applies a further set of requirements including DMARC to anyone who sends more than 5,000 messages per day to Gmail accounts. What that threshold means for a cold sender's actual tier, and what the number is really measuring, is worked through in the Google spam filter.

    Telling them apart from the outside

    A sender cannot read the recipient's headers, so the diagnosis has to be made from the shape of the failure rather than from the verdict itself. Three patterns separate reasonably well.

    Placement that varies by recipient organisation while the message is identical points at a bulk verdict, because the threshold is a per-organisation setting and the content is a constant. Two companies on the same list, same day, same copy, opposite outcomes is the signature.

    Placement that collapses across every recipient at once, shortly after a change to sending domains, mailboxes or authentication, points at the spam side. That verdict travels with the sending identity rather than with the audience, so it does not respect organisational boundaries.

    Placement that degrades gradually over several campaigns from the same domain, with nothing changed in the copy or the configuration, points at an accumulating complaint history. That is a bulk score moving up its scale, and it is the slowest to notice and the slowest to reverse, because the input it responds to is the previous few weeks of recipient behaviour rather than today's message.

    The fourth pattern is the one that misleads. A message landing in a promotions or bulk folder for a recipient who never opens anything from anyone in that folder is not evidence about your sender reputation at all. It is evidence about that mailbox. Reading a single seed inbox as a measurement of the campaign is the most common diagnostic error in this whole area, and it is expensive because it always produces a copy change.

    The repairs, sorted by which verdict they address

    Section illustration: The repairs, sorted by which verdict they address

    The practical value of separating the two verdicts is that it stops a team from applying the wrong repair for six weeks.

    Diagnosing before repairing
    • Yes: Narrow the segment so fewer recipients have no reason to care, which is the only direct input to a complaint rate
    • Yes: Fix authentication and alignment, because an unauthenticated bulk sender is scored as both at once
    • Yes: Remove dead and unverified addresses, since bounces are read as list acquisition rather than as accidents
    • Yes: Make the opt-out mechanism obvious and instant, so a person who does not want this leaves rather than complaining
    • No: Rewrite subject lines to avoid trigger words, which addresses a content model that is not the one filing you
    • No: Increase sending volume to compensate for placement, which raises the pattern signal the score is built on
    • No: Test placement in your own seed inbox and generalise, when the threshold is set per recipient organisation
    Which repair addresses which verdict. The three no rows are the ones teams reach for first and they do not move a bulk score.

    The unsubscribe row deserves a sentence of its own, because it is counter-intuitive to senders who read the opt-out as lost pipeline. A complaint and an unsubscribe are two ways a recipient can express the same thing, and only one of them feeds the score that files your mail. Making the exit easy converts complaints into departures, which is a strictly better trade on the numbers that decide placement. The definitional groundwork on what that rate measures and where it breaks is in spam complaint rate.

    Why the graymail verdict is the one worth aiming at

    There is a version of this article that treats a bulk verdict as a failure. It is not, and pretending otherwise leads to worse decisions.

    A sender scored as bulk with few complaints is in a stable, sustainable place. Microsoft's own scale reserves its bottom band for exactly that sender: bulk, and generating few complaints. The bands above it are earned by complaints rather than by volume. A cold programme cannot stop being bulk, and it can absolutely control which end of that scale it sits at.

    That reframes the objective. The goal is not to look unlike a bulk sender, which is not achievable and produces a lot of superstitious copy editing. The goal is to be a bulk sender that nobody complains about, which is a list problem, a relevance problem and an opt-out problem, in that order.

    Our own practice makes one of those levers structural rather than optional. We send one message per campaign, with no bumps and no thread replies, so a person who does not answer does not hear from us again inside that campaign. That is documented policy rather than a claim about anybody's numbers, and its relevance here is mechanical: a follow-up is delivered to the population that has already declined once, which is the population most likely to complain, and the complaint feeds the score that decides where every other campaign on the domain lands.

    The short version

    Section illustration: The short version

    Graymail and spam are two different verdicts. Graymail is a claim about a sending pattern and a complaint history, and vendors publish it as a numeric score with a threshold the recipient's administrator sets. Spam is a claim about intent, and it is repaired by different things.

    Cold outreach is bulk mail by construction, so the achievable goal is the low end of that scale rather than an exit from it. Complaint rate is the input, and the levers that move it are the segment, the authentication, the list hygiene and how easy the exit is. Subject-line superstition, more volume and seed-inbox testing address none of it.

    If you would rather see a segment narrow enough that the complaint rate takes care of itself, see what a first campaign looks like.

    Questions

    Frequently asked questions.

    Frequently asked questions
    What is the difference between graymail and spam?
    Graymail is bulk mail from a real sender that a filter declines to call abusive, scored on sending pattern and complaint history. Proofpoint's own reference page defines it as bulk email that does not fit the definition of spam because it is solicited, comes from a legitimate source and has varying value to different recipients. Spam is the category where the filter is making a claim about intent.
    Is cold email graymail?
    By the filtering vendors' definitions it sits in that neighbourhood. It is bulk, it comes from a real sender, and its value genuinely varies by recipient. The one place it differs from the newsletter case is consent, since a cold recipient never opted in at all. That gap is what makes complaint rate the binding constraint on the whole motion.
    How do I stop landing in the bulk or promotions bucket?
    Not by rewriting subject lines, because the bulk verdict is scored on complaint history and sending pattern rather than on content. The levers that move it are narrowing the segment so fewer recipients have no reason to care, fixing authentication and alignment, removing dead addresses, and making the opt-out obvious enough that people leave instead of complaining.
    Why does the same message land differently at two companies?
    Because the threshold is set by the recipient's administrator. Microsoft publishes a default bulk complaint level threshold of 7, with 6 under the Standard preset and 5 under the Strict preset, and an organisation picks its policy. Identical copy, same domain, same day, opposite outcomes is the signature of a bulk verdict rather than a spam one.
    deliverabilityspam filterinbox placementcold emailemail infrastructure
    Byline

    About the author.

    Tim Carden

    Tim Carden is CMO / CTO at RevenueFlow, which builds and operates outbound revenue engines for B2B companies. Studied at McGill University.

    Tim Carden · CMO / CTO

    Connect on LinkedIn →
    Your next move

    Ready to scale your outreach?

    We build GTM engines that book real meetings. See the receipts.

    Further reading

    Related articles.

    Cold Email Infrastructure

    Email Warmup Tools: What Actually Moves Placement

    Warmup pools are now free in every major sending tool. What the mechanism actually is, why nobody publishes controlled placement data, and the levers that do work.

    7 min readRead →
    Cold Email Infrastructure

    Instantly.ai Review: What the Plans Include and Who It Suits

    Instantly gives unlimited inboxes on every paid plan and meters contacts and send volume instead. What that changes, where each tier runs out, and who it fits.

    7 min readRead →
    Cold Email Infrastructure

    Migrating Cold Email Tools Without Losing Domain Reputation

    Domain reputation survives a platform switch. Warmup schedules, sending caps and suppression lists do not. How to sequence a migration so mailboxes come through intact.

    7 min readRead →
    Cold Email Infrastructure

    Smartlead Review: Unlimited Mailboxes and What the Plans Cap

    Smartlead includes unlimited mailboxes on every plan and bundles verification on the upper tiers. What rotation does, what it cannot do, and where the ceilings bite.

    7 min readRead →
    Cold Email Infrastructure

    Blacklist IP Search: Whose IP the Lookup Is Checking

    The address your browser reports is almost never the one a receiving server refused. What a blocklist lookup queries, and which IP to run it against.

    8 min readRead →
    Cold Email Infrastructure

    Email Bouncer: What Bouncer Costs and the Product Split

    Bouncer sells verification as credits and deliverability testing as a subscription. Here is the full credit ladder, and the product split that catches buyers out.

    7 min readRead →