Why Most Lead Gen Agencies Fail (And What to Look for Instead)
- Keith Mortier
- Aug 5
- 5 min read
Most B2B lead gen agencies don't fail because they can't find leads. They fail because they get paid for activity and you need outcomes.
Emails sent. Leads delivered. Meetings sourced. Those get measured in the agency's dashboard. Revenue that closes gets measured in your CRM. Almost nobody reconciles the two, and the gap between them is where the budget goes.
Here's the thing. A lead delivered is not a lead worked. That one distinction explains most of what goes wrong below.
Is it the leads, or what happens after the lead arrives?
Usually the second one, and it isn't close.
The most common failure in B2B lead generation isn't a bad list. It's the handoff cliff. The agency delivers a lead, considers the job done, and nobody owns the next 48 hours. The lead sits. It cools. It still gets counted as delivered.
We were brought in to audit a client's existing outbound program, one they were running before we arrived. Going through their CRM, we found 26 prospects who had explicitly asked for something. A guide, a sample report, a pricing sheet. Not one of them ever got it. Twenty-six people raised their hand and got silence back.

That is not a lead generation problem. Every one of those was a generated lead. The sending worked. The follow-through didn't exist, and neither side was measuring it, because the agency's report ended at "delivered" and the client's started at "opportunity created."
Bottom line: ask who works the reply, by name, before you ask anything about volume.
Why do the numbers look fine when the pipeline doesn't?
Because CRM numbers are self-reported and nobody audits them.
Same engagement. The CRM reported 30 open opportunities. We went through them record by record. 18 were real. The rest were duplicates, stale records that should have been closed months earlier, and template defaults nobody had edited. The reported number was inflated by 40 percent.

The wrong number isn't the damage. The damage is that every decision downstream got made against the 30. Budget. Headcount. Forecast. Whether the agency was working.
This happens quietly because CRM hygiene is nobody's job. One number worth knowing: manual data entry carries a 3.7 percent error rate, 260 errors across 6,930 entries, per a 2019 JAMIA study. Small on any single record. It compounds across a pipeline nobody reconciles.
My read: audit the CRM before you buy more leads. If the baseline is wrong you cannot tell whether anything you add is working, and you'll be paying an agency to move a number you can't trust.
Is more volume the fix?
Almost never, and this is where most agency relationships actually die.
When results come back thin the instinct is to send more. But volume before message-market fit just multiplies a message that wasn't landing. It also burns the one asset the whole program sits on, which is your sending domains and your list.
There's a specific trap in the math. Split a modest send across a stack of message variants and no single variant gets enough volume to read a result. Say you split 490 leads across 22 first-step variants. That's about 22 leads per variant. At the low-single-digit positive-reply rates that are normal in B2B cold outreach, 22 sends returns zero or one reply. Zero and one aren't a signal. They're noise, and they look exactly like a winner and a loser.

We have a working internal standard for the minimum sends per variant before we'll read a positive-reply rate as directional, and 22 is nowhere near it. Below that threshold you aren't testing. You're generating confident-looking numbers about nothing and then optimizing toward them.
Bottom line: an agency that answers weak results by proposing more volume, without first checking whether any variant has hit a readable sample, is guessing with your money.
Why does all of this feel so busy for so little output?
Because the work is real even when the results aren't. That's the trap.
Outbound sprawls across tools. The sending platform, the CRM, the enrichment tool, the
scheduler, the inbox, and the spreadsheet nobody admits exists. Harvard Business Review studied 137 workers across 20 teams at three Fortune 500 companies and found people toggle between applications roughly 1,200 times a day. That adds up to just under four hours a week spent reorienting, about 9 percent of working time (Murty, Dadlani & Das, HBR, 2022).
The cost isn't only speed. A CHI 2008 study of 48 workers found interrupted work actually finished faster, 20.3 minutes versus 22.8, but at meaningfully higher stress, 9.46 versus 6.92 on a 20-point scale. People compensate for fragmentation by working harder, not longer. That never shows up in a report. It shows up in turnover.
So the program looks busy, the team feels stretched, the dashboard fills with activity, and the pipeline doesn't move. Activity is not progress. Most agency reporting is built to show you the first one.
What should you actually ask before you sign?
Five questions. Any agency worth hiring answers all five without hedging.
Who works the reply, by name?
Not "our team." A person, with a response-time commitment. If the answer is you, fine, but then it's your program they're feeding and the scope should say so.
What's your written definition of a qualified lead?
Get it in the contract. Most disputes six months in are two definitions colliding.
Will you audit my CRM before we start, and will you show me the audit?
If they won't baseline it they can't prove they improved anything. Neither can you.
How many sends per variant before you read a result?
If they can't answer with a number they aren't testing. They're sending.
What happens to a lead that asks for something?
Make them walk the exact path. This one surfaces the handoff cliff faster than anything else on the list.
The pattern underneath all of it
Every failure above is the same defect wearing different clothes. A metric that reports success while the thing underneath isn't working.
Leads delivered that nobody worked. Opportunities open that closed months ago. Variants tested at unreadable volume. A team busy and going nowhere.
My take: the fix isn't more sophistication. It's insisting every number in the report
traces to something you could check by hand, then actually checking a few at random, early, while the relationship is young enough to correct.
Related: B2B Lead Generation. How we structure outbound so the handoff, the CRM baseline, and the test volume are defined before the first email goes out.
Want your current program audited? We'll baseline the CRM and tell you what's real before recommending anything.
FAQ
How long before B2B lead generation shows results?
First replies land within days. A readable result takes longer than most timelines assume, because it depends on volume per message variant, not weeks on the calendar. Treat any confident verdict inside the first month with suspicion.
Is cold email still effective, and is it legal?
Yes to both, with conditions. In the US, CAN-SPAM requires accurate header and sender information, a working unsubscribe, and prompt handling of opt-outs. GDPR and CASL are stricter and change what's permissible. Effectiveness now depends far more on list quality and relevance than on volume.
Should we build lead generation in-house or hire an agency?
The honest test is whether someone owns replies within hours. If you have that person, an agency can feed them. If you don't, hiring one mostly buys you a bigger pile of unworked leads.
What's a realistic reply rate for B2B cold outreach?
Positive replies in the low single digits are normal for a well-targeted campaign. Anyone quoting a dramatically higher number is usually quoting an open rate, a peak from one narrow segment, or a figure with no denominator attached. Ask for the median, not the best result.
How do I know if my CRM numbers are real?
Pull 20 open opportunities at random and check each by hand. Last real contact, whether the amount was ever edited off the template default, whether the record duplicates another one. If more than a couple don't survive, your reported pipeline describes your data entry habits, not your business.



