Sales Data: The Classes, and Which Ones You Can Buy
Sales data is every record a revenue team holds about its market and about its own selling. It splits into external data bought from a provider, meaning company attributes, contacts, technology and intent signals, and internal data generated by selling, meaning sends, replies, meetings and closed outcomes. Only the first is purchasable, and only the second is unambiguously yours.
Key takeaways
- External sales data is purchasable and decays; internal sales data cannot be bought and does not decay, it goes missing.
- Coverage and accuracy are two different numbers sold under one word, and only coverage is cheap to compute.
- The common shortage is not too few rows but the wrong fields, so write the segment query before buying anything.
- Outcome data only exists if the sending platform writes replies, bounces and opt-outs back to the record.
Sales data is every record a revenue team holds about the market it sells into and about its own selling activity, from company attributes bought from a provider to the outcome of a message somebody sent last Tuesday. The word covers two collections with almost nothing in common except the name, and confusing them is the reason so many teams buy more of one while the shortage sits in the other.
The first collection describes the world outside the company: which companies exist, who works at them, what they run, what they appear to be researching. The second describes what the company itself did and what happened: sends, replies, meetings, opportunities, closes. The first is purchasable and decays. The second cannot be bought at any price and is the only kind that is unambiguously yours.
The two halves, and why only one of them is for sale
External sales data is what a provider sells. It is assembled from public sources, from contributed data, from web crawls and from panels, and it arrives as attributes attached to a company or a person. Firmographic data is the company layer: headcount band, revenue band, industry, location, corporate structure. Technographic data is what the company runs. Contact data is the person layer: name, title, seniority, work address, sometimes a direct dial. Intent data is the behavioural layer, and it is the one with the widest gap between what it is sold as and what it establishes.
Internal sales data is generated by the act of selling and exists nowhere else. It divides again, and the division matters more than the vendor category names suggest. Activity data records what your side did: messages sent, calls placed, meetings held. Outcome data records what the other side did: replied, booked, bought, churned. Both live in the CRM if the write-back loop was built, and in a sending tool's own reporting if it was not, which is the state most programmes are actually in.
- Company attributes, technology, contacts, intent signals
- Priced per record, per credit or per seat
- Coverage and accuracy are two different numbers, and only one is quoted
- Decays continuously from the day it is delivered
- Every competitor targeting your market can buy the same rows
- Sends, replies, meetings, opportunities, closed outcomes
- Costs nothing extra once the campaign is running
- Accuracy is a function of whether the systems write back to each other
- Does not decay, though the account it describes does
- Exists nowhere else and cannot be bought
The asymmetry has a practical consequence that shows up in budgets. External data is the line item, so it gets the attention, the comparison spreadsheet and the renewal argument. Internal data is free and therefore invisible, so nobody notices when a portion of it is being discarded because a sending platform never wrote replies back to the record.
What each class is actually good for
Sales data does three jobs, and a given class is usually good at exactly one of them.
Choosing who to contact. This is a firmographic and technographic job, and it is the one external data does best. The constraint is not the provider, it is your own segment definition: a provider fills the fields it happens to hold, and a segment query filters on the fields your ideal customer profile is written in. Where those two sets barely overlap, buying more rows changes nothing, and the arithmetic behind that mismatch is set out in CRM enrichment.
Deciding when to contact. This is the intent job, and the honest reading of it is narrower than the category name. A signal establishes that somebody at a company did something observable. It does not establish that the person you are about to write to knows about it, or that a budget exists. Where third-party signals are thin, first-party contact produces the same evidence more cheaply, which is why the two are complements rather than substitutes.
Working out what happened. This is entirely internal, and it is the class that decides whether anything gets better. Reply rate by segment, meeting rate by premise, the share of a list that bounced: none of that arrives from a provider, and none of it exists unless the sending platform and the CRM agree about which record a reply belongs to. The write-back loop that makes that possible is the last section of CRM setup for an outbound team, and it is the part most setups skip.
A fourth job gets attributed to sales data and belongs to something else. Forecasting reads pipeline records, and pipeline records are claims a seller typed rather than observations, so a forecast built on them inherits the seller's optimism rather than the data's accuracy. That distinction is worked through in pipeline management in a CRM.
Where the textbook definition misleads
Sales data gets presented as one substance in varying quantities, so the improvement path reads as get more of it. Three properties break that reading.
Coverage and accuracy are different numbers and only one is cheap to compute. A provider can report the share of submitted rows that came back with a value in seconds. Whether the returned value is true today requires somebody to check it against a current source, one row at a time. Those two numbers are reported under one word at purchase time, and the difference between them is the whole of what you are buying. The measurement that governs the economics is match rate, computed on your own table rather than on the provider's sample.
Data you were sold and data you generated fail differently. External data degrades quietly as people change jobs and companies restructure, which is the subject of data decay. Internal data does not degrade, it goes missing: a reply that never reached the record is not stale, it is absent, and nothing on the record announces the gap. Stale external data announces itself eventually, when a message bounces. Missing internal data never announces itself at all.
More rows is the wrong axis for the common shortage. A team that cannot build a campaign usually has plenty of companies and no way to tell which of them fits, because the fields the segment filters on are not the fields anybody bought. The correction is to write the segment query first and let it name the fields, which is nearly always four rather than forty. What to test before paying for a B2B database covers the sample-first version of the same discipline.
- Step 1Define the segment
Write the query first, then list the fields it needs. Four fields, not forty.
- Step 2Buy only those fields
Firmographic and technographic attributes that the query actually filters on.
- Step 3Verify before sending
Addresses checked, catch-all domains handled, duplicates resolved against your own records.
- Step 4Send and capture
Sends, replies, bounces and opt-outs written back to the record with the right owner and timestamp.
- Step 5Read the outcome
Reply and meeting rates by segment and by premise, which is the only evidence that redefines the segment.
How it is used in an outbound programme

The practical test of a sales data set is whether a campaign can be built from it without a research project attached, and that test has four parts.
The list has to be enumerable. A segment you can describe in one sentence and count is a segment; a segment defined by a quality nobody holds as a field is a wish. Building the countable version is the subject of lead list.
The addresses have to survive verification. Coverage on a purchased file is quoted against records returned, not against records that accept mail, and the gap between those two is where a sending reputation goes. Data hygiene is the maintenance side of the same problem.
The suppression side has to be data rather than a setting in a sending tool. A person who asked not to be contacted asked the company, and that fact belongs on the record where every future campaign reads it.
And the outcome has to come back. We run one message per campaign with no bumps and no thread replies, and a later approach is a new campaign on a genuinely different premise. That constraint is documented practice rather than a claim about results, and it changes what the data has to carry: each campaign is a clean test of one premise against one population, so the reply and meeting rates attach to a premise rather than to a blur of five touches nobody can separate afterwards. A programme running sequences cannot tell you which message worked, whatever its reporting shows.
Where third-party intent is the only evidence available, the cheapest counter-evidence is a small campaign. Contact a defined slice, read what comes back, and treat the result as first-party signal on a market where a provider had none. B2B intent data covers what the purchased version predicts and where it stops.
What to do with it
Separate the two collections in your own head before the next purchase conversation, and ask which half the shortage is in. If campaigns cannot be built, the shortage is external and specific: name the missing fields. If campaigns run and nobody can say what worked, the shortage is internal and the fix is a write-back loop rather than a contract.
Buy against a sample you checked by hand. A hundred rows verified against public sources in an afternoon is the only accuracy figure you will ever own, and it beats every published number because it was computed on your market rather than on the provider's.
Keep provider values and human values in separate fields. A rep who spoke to somebody knows something no crawl does, and a bulk enrichment run that overwrites it has destroyed the more valuable of the two.
Report on outcome data by segment and by premise, not in aggregate. An aggregate reply rate across four populations is a number that cannot be acted on, and the segments underneath it usually differ by more than the average suggests.
Related terms and guides
The four external classes each have their own entry: firmographic data for company attributes, technographic data for what a company runs, intent data for behavioural signals, and data decay for why any of them stops being true. Match rate governs what a provider actually costs per usable row, and record matching decides whether an enriched value reaches the right record at all.
On the internal side, CRM setup for an outbound team covers the object model and the write-back loop that decides whether outcome data exists, and pipeline management in a CRM covers what the pipeline records are and are not claiming. For the purchasing decision itself, what to test before paying for a B2B database is the sample-first method, and B2B intent data covers the signal layer specifically.
If the useful next step is producing outcome data on a segment you defined rather than buying more rows about it, see what a first campaign produces.
Frequently asked questions.
Frequently asked questions- What is the difference between sales data and B2B data?
- B2B data usually means the external half only: company attributes, contacts, technology and intent signals assembled by a provider and sold. Sales data covers that plus everything your own selling produces, meaning sends, replies, meetings, opportunities and closed outcomes. The second half is the part no provider can supply and the part most teams under-collect.
- Can you buy sales data, or do you have to generate it?
- Both, and the split is not optional. Company attributes, technology signals, contact records and third-party intent are all for sale. Activity and outcome data are produced only by running campaigns and writing the results back to your records. A team that buys the first and never captures the second can build lists but cannot tell which of them worked.
- How do you know whether purchased sales data is any good?
- Verify a sample by hand rather than reading a published figure. Submit a hundred rows, check the returned values against public sources yourself, and use that result. It takes an afternoon and it is the only accuracy measurement computed on your market rather than on the provider's sample. Repeat it once the contract is live, because coverage and accuracy drift separately.
- Which sales data does an outbound campaign actually need?
- Far less than most teams buy. Write the segment query first and list the fields it filters on, which is usually four or five rather than forty. Add verified addresses and a suppression field, and the list is complete. Everything beyond that is worth having only when a named decision depends on it.