Lead Generation

    The Last Outbound Agency Did Not Deliver: Locating the Failure Before You Buy Again

    The leads were bad names a symptom and no stage. Four places an outbound engagement fails, the counts that separate them, and what to change in the next purchase.

    Editorial illustration for The Last Outbound Agency Did Not Deliver
    August 27, 2026Updated August 28, 20268 min read
    Share:
    The short answer

    An outbound engagement fails at one of four stages: the list, the message, the delivery, or the definition of a qualified meeting. Locate it by comparing five counts, contacted, replied, booked, held and accepted, because each failure collapses the funnel at a different step and each needs a different fix.

    Key takeaways

    • The leads were bad names a symptom rather than a stage, and the four possible causes need four different fixes.
    • Five counts locate the failure: contacted, replied, booked, held, and accepted by sales.
    • Three inputs are the buyer's own: the suppression list, the offer, and how fast replies are answered.
    • Ask the outgoing vendor for the list, the domains, the copy, the replies and the rejection reasons before access ends.

    Reviewed and updated August 28, 2026

    Take an invented but ordinary engagement, with figures chosen to show the shape rather than measured from anybody's account: six months of retainer, forty booked meetings, twenty six of them held, and a sales team that says roughly four were worth attending. The engagement ends, the internal summary says the leads were bad, and that sentence is where the diagnosis usually stops.

    The leads were bad names a symptom and no stage. An outbound programme has four places it can fail, they fail in different ways, and each one leaves different evidence behind. Working out which one you had is the only thing that makes the next purchase different from the last, because the fixes do not overlap: a targeting failure and a definition failure both present as bad meetings and nothing you would change to fix one improves the other.

    The failure has a location

    Four stages, and every disappointing engagement is mostly one of them.

    The list. Wrong companies, or right companies and wrong people. Everything downstream is applied precisely to the wrong audience.

    The message. Right people, and the reason for writing does not land. The list is fine and nobody answers.

    The delivery. Right people, decent message, and a large share of it never arrives. Authentication, domain reputation, sending volume and inbox placement all sit here.

    The definition. Meetings happen and the sales team does not want them. Nothing upstream is broken; the criteria for a billable meeting were never specific enough to exclude the ones you did not want.

    The reason the distinction earns its keep is that the four produce different numbers, so you can tell them apart from data the outgoing vendor already has.

    Targeting failedThe list was wrong
    • Replies arrive at a normal rate
    • The people who reply are the wrong seniority or the wrong company
    • Meetings get held and go nowhere
    • Sales can name what was wrong about each attendee
    • Changing the copy does nothing
    Delivery failedIt did not arrive
    • Reply rate is very low across every version of the copy
    • Bounce rate is high, or suspiciously low with no replies
    • Performance differs sharply between recipient domains
    • Nothing improved when the message was rewritten
    • The vendor talks about volume rather than placement
    The definition failedIt arrived and it counted
    • Meetings were booked and held at a reasonable rate
    • Sales rejected many of them after the fact
    • There was no written criteria list, or it was vague
    • Disputes happened at invoice time
    • Both sides believe they were right
    Three failure signatures that all get reported as bad leads. The numbers behave differently in each, which is what makes the diagnosis possible after the engagement has ended.

    The numbers that locate it

    Five counts, in order, and the stage where the drop is anomalous is the stage that failed.

    Contacted. Replied. Booked. Held. Accepted by sales.

    Take a deliberately invented set of figures, chosen to show the method rather than measured from any engagement: two thousand contacted, sixty replies, forty booked, twenty six held, four accepted. That shape has a normal-looking top and a collapse at the last step, which points at the definition rather than at the list or the message. Reverse it, so that two thousand contacted produces eight replies and three meetings, and the collapse is at the first step, which points at delivery or at a list that does not contain your buyer at all.

    The point of the five counts is that they separate stages that all end in the same complaint. A programme that books plenty and holds few has a scheduling and confirmation problem. A programme that holds plenty and accepts few has a criteria problem. A programme that never gets replies has a problem before any of that.

    Contacted2,000

    The denominator every other number needs and the one most reports omit

    Replied60

    Human replies only. Out of office and auto responses are not replies

    Booked40

    A calendar event exists

    Held26

    The gap from booked is a confirmation and reminder problem, not a targeting one

    Accepted by sales4

    The collapse is here, which points at the criteria rather than at the list

    An invented worked example, chosen to show what a definition failure looks like in the counts. These figures are illustrative and are not measured from any engagement.

    Ask the outgoing vendor for the artefacts

    Section illustration: Ask the outgoing vendor for the artefacts

    The evidence is mostly in their systems, and the moment to ask for it is before the relationship ends rather than after. None of this is unreasonable to request and a vendor's willingness to hand it over is itself informative.

    The list they actually worked, not the list they proposed. The sending domains used, and whether the infrastructure was shared with other clients. The copy that ran, in the versions it ran in. The replies, including the negative ones, because the wording of a rejection frequently names the targeting error precisely. The written meeting criteria, if one existed. And the rejections, with reasons, since a set of rejected meetings with reasons attached is a diagnosis somebody has already half written for you.

    The post mortem artefacts
    • Yes: The list as worked, with the filters that produced it
    • Yes: The sending domains, and whether they were shared with other clients
    • Yes: The copy in the versions that actually sent
    • Yes: Every reply, including the negative ones, in full text
    • Yes: The written qualification criteria, and the rejections with reasons
    • No: A dashboard screenshot of activity totals
    • Depends: Who owned the suppression list, the offer and the reply handling
    What to collect from an engagement that did not work, before the access ends. The last two rows are the ones that are hardest to get afterwards and most useful to have.

    Three failures that are not the vendor's

    This part is uncomfortable and it is the part that decides whether the next engagement goes better, because a buyer who mis-assigns the cause repeats it.

    The suppression list. Customers, live opportunities, partners and protected accounts have to be excluded before the first send, and only you hold that information. Where it arrived late or incomplete, some of the damage attributed to the vendor was yours.

    The offer. A vendor can sharpen wording and cannot decide what you are asking a prospect to do or what they get for it. An engagement that stalled while the proposition was still being argued internally stalled on your side.

    Reply speed. Where replies came to your team rather than the vendor's, the interesting question is how long they sat. A reply answered four days later is frequently a meeting that does not happen, and no vendor performance compensates for it.

    The three inputs no vendor can supply is the pre-purchase version of this list. Reading it after a failed engagement rather than before is a cheaper way to find out which of the three was unowned.

    What to change, by diagnosis

    Section illustration: What to change, by diagnosis

    Each cause points at a different next purchase, which is the practical reason to do the diagnosis at all.

    A targeting failure means the next engagement starts with the account list rather than with the copy, and it means asking each candidate vendor to show the filters they would use and the count those filters return before anything is signed. It may also mean the profile itself is the problem, in which case buying the same service again fixes nothing. Building a profile that changes the target list is upstream of any supplier.

    A message failure means the next vendor's copy process matters more than their volume, so ask who writes it, how many versions run, and what they do when a version underperforms. Treat more sends as an answer to a relevance problem with suspicion.

    A delivery failure means the infrastructure questions move to the front. Whose domains, shared or dedicated, what the warm up looks like, and what happens when a neighbour on shared infrastructure causes a problem. What the architecture has to hold at volume covers what to listen for in the answer.

    A definition failure means the criteria are the negotiation, not the price. Agree in writing what makes a meeting billable, what the rejection window is, and which reasons are valid, before any number is discussed. That is our own position rather than a neutral standard, and we hold it because the alternative turns every invoice into an argument: a meeting counts when the company is in the agreed audience, the person has responsibility for or influence over the relevant area, they agree to a relevant conversation, they attend, and they were not disclosed as an existing customer or live opportunity beforehand. Budget, timing and authority stay outside the billing conditions, since none of them is knowable before the conversation happens. Pinning down what qualified means is the whole of that work.

    Do the vendors deliver at all

    Worth answering plainly, because it is the question underneath most of these post mortems and it deserves better than a defence of the category.

    Outbound run well produces conversations, and the variance between suppliers is wide enough that a single bad experience is genuinely uninformative about the next one. What it is informative about is your own inputs, since those did not change when the supplier did.

    A buyer who has worked with more than one agency is in a stronger position than that, and usually does not feel it. Two engagements against roughly the same market are a comparison, and the interesting question is whether they failed at the same stage. Two failures at the same stage point at an input that did not change between them, which is almost always one of the three above. Two failures at different stages point at supplier selection, and the diagnosis then tells you which of the four questions to press hardest on the next call.

    The honest read on the category is that a disappointing engagement almost always has a locatable cause in one of the four stages, and that some share of those causes sits on the buyer's side of the line. We are not going to put a figure on that share, because we do not have one that would survive being checked. What can be said without a number is that the three buyer-side inputs above are unowned often enough to be worth checking first, and that they are the cheapest thing to fix.

    A buyer who has done this diagnosis is a materially better client for the next vendor, and is also much harder to sell to badly. Both follow from the same thing, which is having a specific account of what went wrong rather than a general disappointment. A vendor hearing the specific version has to answer it; a vendor hearing the general version can answer with a case study.

    Whether to buy again at all

    Section illustration: Whether to buy again at all

    Three conditions that argue for doing it in house instead. The addressable market is small enough that a person can work it by hand and the value per account is high enough to justify that. The offer is still moving, so any external motion is testing a proposition that changes underneath it. Or the constraint was never lead volume, and the meetings you did get were not being worked well after they were booked.

    Where none of those holds, the comparison is arithmetic rather than sentiment, and the cost math on your own numbers is the version to run before shortlisting anybody.

    The short version

    The leads were bad is a symptom with four possible causes, and they need different fixes. Locate it with five counts, contacted, replied, booked, held and accepted, because each failure collapses the funnel at a different step. Collect the list, the domains, the copy, the replies and the rejection reasons from the outgoing vendor while you still have access. Check the three inputs that were yours, suppression, offer and reply speed, before assigning the cause. Then buy against the diagnosis: targeting failures start with the list, delivery failures start with the infrastructure, and definition failures start with the written criteria before any price is discussed.

    If the next arrangement should be per qualified meeting with the criteria agreed in writing before anything sends, you can see what a campaign would look like for your market.

    Questions

    Frequently asked questions.

    Frequently asked questions
    Why did our lead generation agency fail to deliver quality meetings?
    Almost always one of four stages. The list contained the wrong companies or the wrong people, the message gave nobody a reason to answer, delivery meant a share of it never arrived, or the written definition of a billable meeting was loose enough to count meetings sales did not want. The counts tell you which, because each collapses the funnel at a different step.
    What should we ask the outgoing agency for when an engagement ends?
    The list as actually worked with the filters behind it, the sending domains and whether they were shared with other clients, the copy in the versions that ran, every reply including the negative ones in full text, and the qualification criteria with the rejections and their reasons. Ask before the relationship ends, because access disappears with it.
    Do lead generation agencies actually deliver results?
    Variance between suppliers is wide enough that one bad engagement says very little about the next one. What a bad engagement does say something about is your own inputs, since those did not change when the supplier did. Check the suppression list, the offer and reply speed before assigning the cause, because all three are cheap to fix and commonly unowned.
    How do we stop the same thing happening with the next vendor?
    Buy against the diagnosis rather than against the category. A targeting failure means the next engagement starts with the account list and the count it returns. A delivery failure moves the infrastructure questions to the front. A definition failure means the written criteria and the rejection window are settled before any price is discussed.
    lead generationoutboundvendor selectionb2b salesqualified meetings
    Byline

    About the author.

    Ben Carden

    Ben Carden is CRO at RevenueFlow, which builds and operates outbound revenue engines for B2B companies. Previously at Gartner Enterprise. Studied at London School of Economics.

    Ben Carden · CRO

    Connect on LinkedIn →
    Your next move

    Ready to scale your outreach?

    We build GTM engines that book real meetings. See the receipts.

    Further reading

    Related articles.

    Lead Generation

    Already Contracted With Another Lead Generation Vendor: Running Two Without Colliding

    Two vendors collide because neither can see the other's send queue. How to split by account, hold one suppression list, and run a bake-off that decides something.

    8 min readRead →
    Lead Generation

    Two Agencies, One Target List: Who Owns Which Accounts

    Two outbound suppliers on one market will contact the same companies unless the buyer splits it first. How to cut the list, run the exclusion feed and check the overlap.

    8 min readRead →
    Lead Generation

    HR Lead Generation: The Renewal Clock, and Who Actually Signs

    An HR buyer who agrees with every word still cannot act outside their own renewal window. Which triggers are observable, and who signs.

    7 min readRead →
    Lead Generation

    When Outbound Stops Scaling: What Breaks First as You Add Budget

    Doubling the budget rarely doubles the meetings. The five constraints that bind in order, the signature each one leaves in the numbers, and what actually buys more.

    7 min readRead →
    Lead Generation

    IT Lead Generation: Selling to a Buyer Who Runs Your Playbook

    Technology buyers evaluate outbound for a living, and their purchases carry a security review nobody in your meeting owns. Which triggers are real.

    7 min readRead →
    Lead Generation

    Cold Calling as a Lead Source: the Arithmetic Before the Script

    Whether calling can produce ten meetings a month is arithmetic, not opinion. Four numbers decide it, and working backwards is the version that stops bad hires.

    7 min readRead →