Glossary

    Match Rate: The Number That Decides What a Data Provider Costs You

    The short answer

    Match rate turns a per-credit price into a real one: divide price per credit by match rate and you get cost per usable record, which is what actually orders your providers. Measure it on your own list, name the denominator before believing any figure, and never read coverage as evidence of accuracy.

    Key takeaways

    • Effective cost per usable record is price per credit divided by match rate. A provider matching four rows in ten must be less than half the price of one matching eight to break even.
    • Rows submitted, rows attempted and rows billed are three different denominators, and a published rate is free to use whichever flatters it.
    • A rate computed over an already-filtered set reports on survivors, so a run that dropped almost everything can post a near-perfect match rate. Assert the denominator before reading the rate.
    • Match rate and accuracy are independent. A returned address is a claim, and verification is a separate step with its own cost.

    Match Rate: The Number That Decides What a Data Provider Costs You

    Match rate is the share of records submitted to a data provider for which the provider returns the attribute you asked for. Submit ten thousand rows asking for a work email address, get six thousand addresses back, and the match rate on that run is sixty percent. It is quoted as a percentage of rows submitted, it is the headline number on almost every enrichment vendor's site, and it is the single figure that determines what a provider actually costs you per usable record rather than per credit.

    The term exists because list prices are meaningless without it. Two providers can quote per-credit prices that differ by a factor of three, and the expensive one can be the cheaper purchase, because a credit spent on a row that comes back empty bought you nothing. Before teams had a word for this they compared vendors on price per record and were routinely surprised by the invoice, since the thing being priced was an attempt rather than a result. Match rate is what makes those two prices comparable.

    How match rate actually works

    The arithmetic looks trivial and hides a moving denominator. Rows submitted, rows attempted and rows billed are three different counts, and a vendor is free to compute its published rate over whichever one flatters it.

    Rows submitted is your input file. Rows attempted is the subset the provider actually looked for, which can be smaller if it silently drops records it considers unresolvable, out of region, or missing a required key. Rows billed depends entirely on the pricing model: some providers bill per submission, some bill only on a returned value, and some bill on submission but refund misses as credits. A rate computed over rows attempted will always beat the same run measured over rows submitted, and only the second one describes what happened to your list.

    Rows on your list

    The population you actually need to reach

    Rows submitted

    What survived your own filters and got sent to the provider

    Rows attempted

    The subset the provider looked for; misses here can be invisible

    Rows returned with a value

    The numerator in most published match rates

    Rows that survive verification

    The only count that turns into a message anybody receives

    The five counts behind one match rate, and why the number changes depending on where you measure it.

    There is a second reason the published number rarely survives contact with your data, and it has nothing to do with anyone being dishonest. A vendor's match rate is measured on the vendor's own sample, which tends to be built from the segments where that vendor is strong: a provider whose coverage comes from professional network data reports well on mid-market technology companies in North America and thins out on small local businesses in Europe. Your list is not that sample. The rate against your own list is the only one that means anything, and it is cheap to establish by submitting a few hundred rows before signing anything.

    Chaining several providers rather than betting on one is the standard response, and the ordering logic follows directly from these numbers. The full version of that reasoning is in waterfall enrichment, and the credit models that make the same run cost wildly different amounts are worked through in Clay's pricing model.

    The arithmetic that decides provider order

    Effective cost per usable record is price per credit divided by match rate. That one line is what actually decides which provider you run first, and it is why per-credit prices on their own tell you nothing.

    The following numbers are illustrative, so substitute your own. Suppose one provider matches four rows in ten on your list and a second matches eight in ten. The first needs two and a half credits to produce one usable record, the second needs one and a quarter. For the low-match provider to be the cheaper source, its per-credit price has to be less than half the other's, not merely lower. Most vendor comparisons stop at the per-credit price, which is the number that has been made comparable and the number that decides least.

    Run the same division across every provider you are considering, using match rates you measured on your own sample rather than published ones, and the ordering usually rearranges itself. It also tends to reveal that the cheapest provider belongs first in the chain rather than nowhere: a low-cost source with a modest match rate is excellent as a first pass, because every row it resolves is a row the expensive source never has to be asked about.

    One caution on the division itself. It is only valid when the match rates are measured over the same denominator on the same list. A rate one vendor computed over rows attempted and another computed over rows submitted are not comparable numbers, and dividing prices by them produces a confident ranking of nothing.

    The same division answers a question teams usually settle by instinct, which is how deep the chain should go. Each additional provider in the order is asked only about the rows nobody before it resolved, so its effective cost is computed against a smaller and progressively harder population, and its own match rate on those leftovers is lower than its published rate by definition. There is a point where the next source costs more per usable record than the record is worth to you, and that point is calculable rather than a matter of taste. Working it out once stops the chain growing by habit every time somebody demonstrates a new tool.

    Where the textbook definition breaks

    The clean definition assumes the denominator is stable and known. It usually is neither, and this is the trap that costs the most money.

    A rate computed over an already-filtered set reports on the survivors. If a provider only attempts the rows it believes it can resolve, its match rate is measured on a population it selected for winnability, and it reads close to perfect while the coverage of your actual list is poor. The extreme version is the tell: a run where almost every row was dropped before the attempt can post a near-flawless match rate and deliver almost nothing. Nothing errors, nothing looks wrong, and the reported percentage is arithmetically correct.

    Always assert the denominator before reading the rate. Ask what number sat underneath it, confirm that number is your submitted row count rather than the provider's attempted count, and treat any rate whose denominator you cannot name as an unverified claim. This applies to your own reporting too, not only to vendors: an internal dashboard that computes quality percentages over rows that already passed an upstream filter will report health at exactly the moment a run has failed completely.

    The general shape of this error is worth recognising because it recurs everywhere a percentage is reported by the system being measured. Our own 2026 benchmark report, built on 1,413,405 sends with every bounce-folder message classified by hand, found a platform bounce counter reading 3.04% where genuinely bad addresses were 1.27%, an inflation of 2.4 times, because the counter was counting a broader set of messages than the label implied. That figure measures bounces rather than enrichment, and the analogy is what transfers: a reported rate describes whatever population the reporting system chose, and only a hand count over a denominator you defined tells you what happened. The full methodology is in the 2026 cold email benchmark report.

    The second break is simpler and just as expensive. Match rate and accuracy are independent, and only one of them is easy to measure. A match means the provider found something it believes belongs to that record. It does not mean the value is current, and for email it does not mean the mailbox exists. A returned address is a claim, and turning it into a fact needs a separate verification step, which is covered in email verification tools. Providers with the highest match rates are frequently the ones most willing to return a pattern-guessed address, which is exactly the trade you would expect: loosening the confidence threshold raises coverage and lowers precision at the same time.

    Questions that make a quoted match rate meaningful
    • Yes: The denominator is rows you submitted, not rows the provider chose to attempt
    • Yes: The rate was measured on your list, not the vendor's published sample
    • Yes: You know whether misses are billed, refunded, or free
    • Yes: Returned values were verified separately before being counted as usable
    • Yes: The sample covered your weakest segment, not your easiest one
    • No: The rate is quoted as a single headline figure with no denominator named
    • No: Coverage is being read as evidence that the returned values are correct
    What to establish about a match rate before it is allowed to influence a purchase.

    What to do with it

    Measure it yourself on a sample drawn from the hard part of your list. A few hundred rows from the segment you are least confident about tells you more than a thousand rows from the middle, because the middle is where every provider looks good.

    Then compute effective cost per usable record, not price per credit, and use that to order your providers rather than to pick one. Re-measure when your target segment changes, since a match rate is a property of the pairing between a provider and a population rather than a property of the provider.

    Record the misses as deliberately as the hits. A row that came back empty and a row the provider never attempted look identical in the output file, and only one of them is worth submitting to the next source in the chain. Keeping the rejection reason alongside each row is what lets you tell an exhausted population from an unattempted one, and it stops you paying twice for the same answer a quarter later.

    And keep the verification step separate in your reporting, so that coverage and correctness never collapse into a single number on a dashboard. The moment they merge, the incentive quietly shifts toward whichever source returns the most values, which is not the same thing as whichever source is right.

    CRM enrichment is the workflow that consumes this number, and record matching decides what counts as a match in the first place. Both have their own entries here. Data decay explains why a match that was correct at the time stops being correct later.

    On the tooling side, best email finder tools compares the sources whose match rates you would be measuring, and Apollo's email finder is a useful single-vendor worked example of coverage and confidence being traded against each other. For the chain rather than the individual source, waterfall enrichment is the ordering argument, and email verification tools covers the step that turns a returned address into a usable one.

    If you would rather see a campaign run against a list where the usable-record count was established before anything was sent, see what a first campaign looks like.

    Questions

    Frequently asked questions.

    Frequently asked questions
    What is a good match rate?
    There is no portable answer, because a match rate is a property of the pairing between a provider and a population rather than of the provider. The same source can perform strongly on mid-market technology companies and poorly on small local businesses. The only useful figure is the one you measure on a sample drawn from your own list, including its hardest segment.
    Why does a vendor's published match rate rarely hold on my own list?
    Published rates are measured on the vendor's own sample, which is usually built from the segments where its coverage is strongest. That is ordinary selection rather than dishonesty, and it means the number describes the sample rather than your list. Submitting a few hundred of your own rows before signing replaces the claim with a measurement.
    Does a match mean the email address works?
    No. A match means the provider found a value it believes belongs to that record. Whether the mailbox exists is a separate question answered by verification, and providers with the highest coverage are often the most willing to return a pattern-guessed address. Loosening the confidence threshold raises match rate and lowers precision at the same time.
    How does match rate decide provider order in a chain?
    Each source in the chain is asked only about rows nobody before it resolved, so it works a smaller and harder population and performs below its published rate. Ranking sources by cost per usable record, then running the cheapest first, means every row an inexpensive source resolves is a row the expensive one is never asked about.