Glossary

    Lead Scoring: What a Single Number Buys You, and What It Destroys

    The short answer

    Lead scoring is the practice of assigning a numeric value to each lead so a sales team can work them in order. Points come from fit attributes such as industry and seniority, and from behaviour such as page visits and downloads. Above a threshold the lead is routed to a salesperson; below it, it waits.

    Key takeaways

    • Fit and intent behave completely differently, and summing them into one score destroys the distinction that decides what to do.
    • A score is a ranking, not a probability, so gaps between scores are not comparable and an average score means nothing.
    • The routing threshold is normally set by team capacity, which means qualified is defined by staffing rather than by the buyer.
    • Behaviour scoring can only rank an audience, never a market, so it is silent about the accounts you most want to reach.

    Lead Scoring: What a Single Number Buys You, and What It Destroys

    Lead scoring is the practice of assigning a numeric value to each lead so that a sales team can work them in order. Points are awarded for attributes of the company or person, for behaviour such as page visits and content downloads, and sometimes deducted for disqualifying signals. Above a threshold, the lead is routed to a salesperson. Below it, it waits.

    The model is a prioritisation device and nothing more. It answers "who should be called first" with a number, and every serious problem with lead scoring comes from asking it a question it was never built to answer.

    What goes into a score

    Scores are built from two ingredients that behave very differently.

    Fit signals describe what the company and the person are, independent of anything they have done. Industry, headcount, revenue band, technology in use, seniority, geography, whether they match the ideal customer profile. Fit is stable, knowable before any interaction, and available for every company in a market rather than only for the ones who visited your site.

    Intent signals describe what they have done. Pages viewed, an assessment completed, a pricing page opened twice, a webinar attended, an email opened. Intent is volatile, arrives only from people already in contact with you, and decays quickly.

    A single score adds these together, and the addition is where the information is lost. A perfect-fit company whose intern downloaded one PDF and a poor-fit company whose CEO read the pricing page four times can arrive at the same number, and that number tells a salesperson to treat them identically. The full version of this argument, with what to do instead, is in qualified lead marketing.

    FitWhat they are
    • Industry, size, seniority, technology, geography
    • Stable for months or years
    • Knowable for every company, contacted or not
    • Wrong fit is usually terminal for the deal
    IntentWhat they did
    • Pages viewed, forms filled, events attended
    • Decays in days
    • Only exists for people already interacting with you
    • Weak intent today says little about next quarter
    The single scoreWhat survives the addition
    • One rank order for a queue
    • Neither of the two inputs is recoverable from it
    • Two opposite situations can share a number
    • Useful for ordering work, useless for deciding what to say
    Two ingredients that get summed into one number. They answer different questions and decay at different speeds.

    How a model is usually built

    Most scoring models are built one of two ways, and the difference decides how much to trust them.

    Hand-assigned weights. Someone decides a director title is worth 15 points and a pricing-page visit is worth 20. This is fast, transparent and entirely opinion. It encodes what the team believes rather than what happened, and it tends to overweight whatever behaviour is easiest to track.

    Fitted from outcomes. The weights are derived from historical closed-won and closed-lost data, so the model reflects observed conversion rather than belief. Better, and it requires enough closed deals for the pattern to be real. With a few dozen wins, a fitted model is describing coincidence with great confidence.

    There is a third pattern worth naming because it is common and quietly circular: weights fitted on which leads the sales team accepted. That model learns to predict sales team behaviour, including its biases, and reports the agreement as accuracy.

    Where the textbook definition breaks

    A score is a ranking, not a probability. Ranks are ordinal. The gap between 90 and 80 is not comparable to the gap between 50 and 40, and averaging scores across a cohort produces a number with no meaning. Teams that report "average lead score improved to 62" have measured nothing.

    Thresholds get set by capacity and then interpreted as quality. The bar is usually chosen so the volume above it matches what the team can work. That is a sensible operational decision, and it means "qualified" is defined by staffing rather than by the buyer. When the team grows, the same leads become qualified, and nothing about them changed. The definitional argument this threshold is standing in for is set out in MQL vs SQL.

    Behaviour scoring only sees people who came to you. A perfect-fit company that has never visited your site scores zero on every intent signal, which is correct and useless. Scoring cannot rank a market, only an audience, so it has nothing to say about the accounts you would most like to reach.

    Score inflation is invisible. A model that awards points for common actions accumulates high scores over time as contacts age in the database, and old contacts drift upward past the threshold without any recent interest. Any model that adds without decaying will eventually route stale records to a salesperson.

    It gets used to decide the message, which it cannot do. The score has already destroyed the distinction between fit and intent, so a message written from the score alone is written from an ambiguity. What to say comes from the underlying signals, not from their sum.

    Auditing a lead scoring model
    • Depends: Fit and intent are visible separately, not only as a total
    • Depends: Intent points decay with time rather than accumulating forever
    • Depends: Weights come from closed outcomes, or their opinion basis is acknowledged
    • Depends: The model is not fitted on which leads sales chose to accept
    • Depends: The threshold is documented as a capacity decision, with its current value
    • Depends: Someone checks quarterly whether high scores actually convert better
    Six checks on a scoring model that separate a working prioritisation tool from a comforting number.

    Two leads, one score

    The clearest way to see what the addition costs is to run two records through a plausible model.

    Suppose fit points are awarded for company size, industry match and seniority, and intent points for pricing-page visits, content downloads and email engagement. The threshold for routing to a salesperson is 70.

    Lead A works at a 900-person logistics company squarely inside your profile, holds an operations director title, and has visited once, landing on a blog post from a search. Fit is close to maximal, intent is close to zero. The model returns 72.

    Lead B is a marketing coordinator at an eleven-person agency that could never buy at your price. She has downloaded four pieces of material, opened everything, and viewed the pricing page twice while writing a comparison for a class. Fit is close to zero, intent is close to maximal. The model returns 71.

    Both route. Both arrive in the same queue, in effectively the same position, described by the same number. A salesperson working the queue in order will treat them the same way for the first minute, which is exactly the minute in which the difference is decidable.

    The disposals are opposite and neither is difficult. Lead A is a company you want, with one visit and no reason to be talking to you yet, so the correct action is a specific approach on a premise about their business rather than a call about a blog post they barely read. Lead B is a fast, courteous no.

    Nothing about that judgment required a better model. It required the two components to be visible, and the single score is the only thing that took them away. This is the practical cost of the addition: the model did not make a mistake, it discarded the input that made the decision obvious, and then presented what was left as a judgment.

    The same failure runs at the aggregate level. A dashboard reporting that qualified volume rose 14 percent cannot say whether more good-fit companies appeared or whether existing contacts accumulated clicks, and those imply different actions.

    What to do instead of one number

    The practical fix is small and costs nothing: keep two scores and never add them.

    A fit score and an intent score, reported as a pair, preserve exactly the information a single number destroys, and they produce four groups a team can act on differently. High fit with high intent is the queue. High fit with low intent is the outbound list, and it is usually the largest and most neglected group in any database. Low fit with high intent is where a polite, fast disqualification saves everyone time. Low fit with low intent needs nothing at all.

    That second group is worth dwelling on. It contains every company you would happily sell to that has not yet raised a hand, which is the population outbound exists for, and a single-score model files them alongside genuinely bad leads because both sit under the threshold.

    Keeping the pair costs nothing to implement. Most platforms already compute the components separately before summing them, so the change is usually a reporting decision rather than a modelling one: expose both columns, route on the pair, and stop reporting the total. The one habit that has to change alongside it is the language, because a team accustomed to saying "an 82" has to start saying "high fit, low intent", and that sentence is the whole benefit. It names a situation and implies an action, where a number names neither.

    The same logic applies to any composite metric assembled from components that decay at different rates. Summing them produces a figure that is stable, comparable and unable to answer the question anyone actually has.

    The related discipline of deciding who is worth a salesperson's hour, using evidence rather than points, is lead qualification, and the frameworks used further down the deal are BANT and MEDDIC.

    Lead qualification is the human judgment this automates part of. Inbound lead is what most scoring models are built on top of. And lead nurturing is what usually happens to everything below the threshold.

    The short version

    Lead scoring ranks leads so a team can work them in order. Keep fit and intent apart, decay the behavioural half, fit the weights on outcomes rather than opinions, and remember that the threshold is a staffing decision. Above all, do not let a number that cannot rank a market decide which markets you go after.

    The high-fit, low-intent group is the one a score cannot help with, and it is the one we build campaigns against: see what that list produces.

    Questions

    Frequently asked questions.

    Frequently asked questions
    What is the difference between fit and intent in lead scoring?
    Fit describes what a company and person are: industry, size, seniority, technology in use. It is stable and knowable for every company in a market. Intent describes what they did: pages viewed, forms filled, emails opened. It decays in days and exists only for people already interacting with you.
    Why is a single lead score a problem?
    Because two opposite situations reach the same number. A perfect-fit company with one visit and a poor-fit company with heavy engagement can both score 71, and the score tells a salesperson to treat them identically. The correct actions are opposite, and the information needed to see that was discarded in the addition.
    How should lead scores be built?
    Keep fit and intent as two separate scores and never add them. Decay the intent half so old records cannot drift upward past the threshold. Fit the weights on closed outcomes rather than opinion where you have enough deals, and avoid fitting them on which leads the sales team chose to accept, which is circular.
    What does a low lead score actually tell you?
    Only that the sum was small, which conflates two very different groups. Good-fit companies who have not raised a hand sit below the threshold alongside genuinely poor-fit contacts, and the first group is the population outbound exists for. A single-score model files both as the same thing.