ICP Scoring: The Fit Score That Ranks a Cold Market
Filters decide who belongs in the addressable set. A weighted fit score decides which of the survivors gets worked first, and only when capacity demands it.

ICP scoring turns a customer profile into a weighted number over company attributes, so companies that already passed the profile filters can be ranked against each other. It reads firmographics rather than behaviour, which is why it can rank companies that have never contacted you. Build one only when the addressable set exceeds capacity.
Key takeaways
- Filters decide membership of the addressable set and a score decides order within it, so the two answer different questions and both can be right.
- A fit score is computable for a company that has never heard of you, which is the population an outbound programme works and the one behavioural scoring cannot rank.
- Weights are defensible when fitted on your own closed-won accounts or labelled as opinion, and misleading when hand-assigned and presented as derived.
- An attribute your provider cannot supply for most of the list scores zero for those rows, so unobservable companies sink without being worse fits.
Reviewed and updated September 2, 2026
A target list arrives with 4,000 companies on it and a team that can work 300 a month. Every company on the list passed the ideal customer profile, because the profile was built as a set of filters and 4,000 rows cleared them. The question the filters cannot answer is which 300 to start with, and that question is what an ICP score exists to settle.
ICP scoring is the practice of turning a customer profile into a weighted number, so that companies which pass the profile can still be ranked against each other. It is a different instrument from lead scoring, it fails in different ways, and the most common mistake is building one before the situation calls for it.
Where the score sits, and what it replaces
A profile expressed as filters produces a binary answer. Either a company is in the addressable set or it is not, and this site argues in building an ideal customer profile that a field which cannot be expressed as a filter will not survive the trip into sourcing at all. That argument holds. Filters are what a data provider accepts, and a profile that cannot be entered as filters has not changed the target list.
A score does something a filter cannot: it orders what survives. It takes the same attributes, gives each a weight, and returns a number that sorts the set.
The decision between them is capacity. Where the addressable set is smaller than what the team can work, ranking it is a solution to no problem, and building a scoring model is a way of feeling rigorous while contacting the same companies in a different order. Where the set is several times larger than capacity, the ordering decides which companies are contacted this quarter and which are contacted never, and that is a real decision being made either way. Made deliberately it is a score. Made by accident it is alphabetical order, or whatever order the export happened to arrive in.
- Answers whether a company belongs in the set
- Enterable directly into a data provider
- Auditable line by line
- Says nothing about which of the survivors to start with
- Right when the addressable set is close to capacity
- Answers which survivors to work first
- Requires weights somebody has to justify
- Hides its inputs behind one number
- Only meaningful over companies that already passed the filters
- Right when the set is several times larger than capacity
What goes into the number
An ICP score reads company attributes and nothing else. That constraint is the whole reason the instrument exists, and it is where it separates from lead scoring.
This site's lead scoring entry sets out the general problem with adding fit and behaviour into one figure: two opposite situations arrive at the same number and neither input is recoverable afterwards. It also names the limitation that makes ICP scoring necessary. Behavioural scoring can only see people who came to you, so it ranks an audience rather than a market. Every company you would most like to reach scores zero on every intent signal, correctly and uselessly.
A fit score is computable for a company that has never heard of you, which is the entire population an outbound programme works. The attributes available are the ones a provider or a public source can supply: industry, headcount, revenue band, geography, technology in use, corporate structure, and observable events such as an opening for a role your product supports. This site covers the raw material in firmographic data, and the event layer in intent data, which is worth holding separately for reasons the next section covers.
The temptation once the model exists is to fold behavioural signals in, because they are available for some rows. Resist it for as long as you can. The moment intent joins the score, the score can no longer rank a company nobody has contacted, and the instrument becomes a lead score with an ICP label on it.
Setting the weights, with an invented worked example

The arithmetic below is entirely invented. The company, the segments and every number are illustrative, and none of them describes a real client or a measured result.
Take an invented profile with four attributes and a 100-point scale.
Suppose headcount carries 40 points, awarded in full between 50 and 500 employees and half between 500 and 2,000. Suppose a particular technology in the stack carries 25. Suppose an open role of a specific kind carries 20. Suppose being headquartered in a market you already sell into carries 15.
An invented company with 180 employees, the technology in place, no relevant opening and a home market inside your footprint scores 40 plus 25 plus 0 plus 15, which is 80. A second invented company with 900 employees, no sign of the technology, two relevant openings and the same geography scores 20 plus 0 plus 20 plus 15, which is 55.
The arithmetic is trivial. The interesting part is the argument for those four numbers, and there are only two honest ways to produce them.
Fitted from your own closed-won. Take the accounts that actually bought, and check which attributes they share against the accounts that did not. This is the better method and it has a hard prerequisite: enough closed deals for the pattern to be real. With a few dozen wins, a fitted model is describing coincidence with considerable confidence.
Assigned by hand and labelled as opinion. Somebody decides the numbers. This is fast, transparent and entirely belief, and it is fine as long as the document says so. What is not fine is a hand-assigned model presented as though it were derived, because that is the version nobody revisits.
There is a third pattern worth naming because it looks like the first one. Weights fitted on which accounts the sales team accepted teach the model to predict the sales team, including whatever it was already getting wrong, and then report the agreement as accuracy.
- Step 1Write the profile as filters first
A score can only rank companies that already passed. Scoring the unfiltered market ranks companies you would refuse
- Step 2Count the survivors against capacity
If the set is close to what the team can work, stop here. There is nothing to rank
- Step 3Choose attributes a provider can actually supply
An attribute you cannot source for every row scores zero for most of the list and quietly buries those companies
- Step 4Derive the weights from closed-won, or label them opinion
Both are legitimate. Only one of them is legitimate while being described as the other
- Step 5Band the output and forget the number
Tiers are what a rep can act on. The two-point gap between 71 and 73 means nothing
Bands, not points
A score is a rank order rather than a quantity. The gap between 90 and 80 is not comparable to the gap between 50 and 40, and averaging scores across a cohort produces a figure with no meaning attached to it. A team reporting that its average ICP score improved has measured nothing.
What survives that limitation is banding. Sort the list, cut it into two or three tiers, and treat the tier as the operating unit. The top tier gets the research and the specific message. The middle tier gets the campaign. The bottom tier waits, and it waits without a plan to revisit it unless something about the company changes.
Where the cut falls is a capacity decision rather than a quality one, and it is worth saying so out loud. The bar is usually set so that the volume above it matches what the team can work. That is sensible, and it means the word qualified in this context is defined by staffing. Grow the team and the same companies become qualified without anything about them having changed. This site follows that argument to its conclusion in what one score destroys.
Where the model breaks

Four failures recur, and three of them are invisible from inside the model.
A missing attribute scores as a bad one. Where technology data covers only part of your list, every row with no data scores zero on that attribute and sinks below rows where the data happened to exist. Those companies are not worse fits. They are less observable. Check coverage per attribute before you trust the ranking, and either drop an attribute you cannot source broadly or score its absence as neutral rather than as a miss.
The weights encode what is easy to measure. Headcount is available for everybody and correlates with a great deal, so it tends to end up carrying the model. Whether it deserves to is a question the closed-won data can answer and an opinion cannot.
The score ages. Firmographic attributes are stable for months rather than for ever. Companies grow, change technology and get acquired, and a score computed last year on a list you are working this quarter is describing companies that have moved. Recompute on a schedule and treat the recompute date as part of the record.
The number replaces the reason. This is the expensive one. A rep handed a score of 82 knows the order to work in and knows nothing about why this company is here, which is exactly the material the first line of the message needs. Store the attribute values alongside the score rather than the score alone, for the same reason this site argues for storing a dated action rather than a temperature.
- Yes: The addressable set is several times larger than the team's monthly capacity
- Yes: Every attribute is sourceable for most of the list, and coverage was measured
- Yes: The weights are either fitted on closed-won or labelled as opinion
- Yes: Attribute values are stored beside the score, not replaced by it
- Yes: The output is banded into tiers a rep can act on
- No: Behavioural signals are folded into the same number
- No: Weights were fitted on which leads the sales team accepted
How platform scoring relates to this
Most CRM and marketing platforms ship a scoring feature, and the good ones already keep fit and behaviour apart rather than adding them. HubSpot's split between the two, and which subscription unlocks it, is worked through in HubSpot lead scoring; the two-number model in Salesforce, where fit and behaviour are graded and scored separately, is in Salesforce lead scoring and grading.
The relevant point for outbound is that these tools score records that exist in the CRM. An ICP score for an outbound programme has to be computable for companies that are not in the CRM yet, which usually means it lives in the sourcing layer rather than in the platform, and gets written onto the record when the company enters it.
The short version

An ICP score is a weighted fit number over company attributes, used to rank a target list that has already passed the profile filters. It exists because behavioural scoring can only rank an audience, and an outbound programme needs to rank a market.
Build one only when the addressable set is several times larger than what the team can work. Use attributes a provider can supply for most of the list, measure that coverage before trusting the ranking, and either fit the weights on your own closed-won accounts or label them as opinion.
Read the output as bands rather than points, since the number is ordinal and averaging it means nothing. Keep behaviour out of it, keep the attribute values beside the score so the message has something to be built from, and recompute on a schedule because firmographics move.
If the more pressing problem is that the addressable set is too small to need ranking, the fix is upstream in the profile itself. See what a campaign against a properly built target list looks like in your market.
Frequently asked questions.
Frequently asked questions- What is ICP scoring?
- It is the practice of assigning weights to company attributes so that a target list can be ranked rather than only filtered. The inputs are firmographic and technographic attributes plus observable events, and the output is a number used to sort companies that have already passed the ideal customer profile. It reads what a company is rather than what anyone has done.
- How is ICP scoring different from lead scoring?
- Lead scoring usually adds fit attributes and behavioural signals into one figure, which means it can only rank people who have already interacted with you. An ICP score reads company attributes alone, so it works on companies that have never heard of you. That is the whole reason an outbound programme needs one instead.
- How do I set the weights?
- Two honest methods exist. Fit them against your own closed-won accounts, which needs enough closed deals for the pattern to be real rather than coincidence. Or assign them by hand and record in the document that they are opinion. Weights fitted on which leads the sales team accepted teach the model to predict the team rather than the market.
- Should I use the score as a number or as a band?
- As a band. A score is ordinal, so the gap between 90 and 80 is not comparable to the gap between 50 and 40, and averaging scores across a cohort produces a figure with no meaning. Cut the sorted list into two or three tiers, act on the tier, and store the attribute values beside the score so the message has something to draw on.
About the author.
B2B cold email experts helping companies generate qualified leads through done-for-you outreach campaigns.
RevenueFlow Team
Explore more.
Ready to scale your outreach?
We build GTM engines that book real meetings. See the receipts.
Related articles.
Sales Qualifying Questions, Sorted by What They Test
A numbered list of questions is not a method. Sorted by the five things they establish, the same questions become one, and the weak ones become visible.
Buyer Experience: The Part Outbound Decides First
Most buyer-experience work goes into the demo and the proposal. The chapter that runs first is a message the buyer never asked for, and three properties decide it.
High Pressure Sales Tactics: What They Cost in B2B
Pressure selling manufactures scarcity, consensus and fatigue. The consumer has a statutory cancellation window against it and the B2B buyer has none.
Sales Pitch Conclusions: 12 Lines and What Each Assumes
Pitch advice covers the opening and gives the ending one sentence about recapping. Twelve closing lines, grouped by situation, with the assumption each rests on.
Pay Per Meeting Pricing When the Deal Is Small
Pay per meeting is decided by your deal economics before it is decided by the vendor. The arithmetic that settles it, and what to buy when it does not clear.
12 Sales Pitch One-Liners and the Structure Behind Them
A one-line pitch has three slots and about two seconds. Twelve invented examples by situation, plus the test that separates a pitch from a description.