How to Tell If Your Outbound Agency Is Actually Delivering Meetings
No standards body defines a meeting, so agencies grade their own homework. A seven-step audit you can run on your own calendar and CRM data before renewal.
From 24 named sources. Claims are tagged [Verified] (primary, sample disclosed), [Secondhand] (traceable, no sample), [Unsourced] (chain breaks at a vendor blog), [Contradicted] or [No evidence]. Sources and gaps: the Sources section at the foot of this page.
Last updated: 31 August 2026
"Lied to on our sales calls about having a full time person."
That is a buyer on r/b2bmarketing, writing after signing. On X, another put the same feeling in numbers: "they won't tell you their close rate is 4% until month three."
If you are mid contract and suspect the monthly report is managing you rather than informing you, the problem is structural, not paranoid. Every unit of measurement in outbound is self-defined, and the party being measured picks the definition. No regulator or standards body defines a "meeting" or a "qualified meeting" for agency work, so an agency reporting 20 meetings and one reporting 4 can be delivering identical value.
The audit cannot be run on their report. It has to run on your data.
Of the nine most widely circulated outbound benchmarks I traced back to source, six break their citation chain at a vendor blog with no disclosed sample, two are misattributed to studies that do not contain them, and one is a rounded real figure.
Value leaks at every rung, and agencies invoice near the top.
| Rung | Defined by | What it means |
|---|---|---|
| Booked | Agency, no standard | Instantly: meetings over prospects emailed, roughly 1% of sends, no sample [Unsourced] |
| Confirmed | Calendar | An accepted invite. Ivris: "an accepted calendar invite is not attendance" |
| Held | Observable | The external party joined. Calendly and RevenueHero allow a no-show mark only after the scheduled start, so it is timestamped and checkable |
| Qualified | Most contested | ICP match, or BANT, or "answered the phone." Leadium: "A booked meeting is not a qualified meeting" |
| SAL / SQL | SiriusDecisions | Sales validates agreed criteria, then prioritises on timeframe, budget, role [Verified] |
| Opportunity | CRM | A record carrying a dollar value and a close date [Verified] |
| Closed-won | Finance | Revenue |
The only near-convention is the SiriusDecisions (now Forrester) Demand Waterfall, created 2006 and re-architected 2012, formalising Inquiry, MQL, SAL, SQL, Close [Verified]. It defines marketing-to-sales handoff stages, not meeting gradations.
Most agencies invoice at "booked" or self-defined "qualified," not at held. Belkins documents pay-per-appointment at $25 to $100, and pay-per-qualified-lead reaching several hundred dollars [Secondhand]. DemandNexus, billing only for BANT-qualified meetings actually attended, is the exception. You pay before the value event.
| Transition | Figure | Source and sample |
|---|---|---|
| Booked to held | 80% hold, 20% no-show | Operatix via Gradient Works, no sample [Unsourced] |
| Cold-booked to held, by lead time | 6.9% same day, 9.6% next day, 23.0% at 8+ days | Reply.io, 2,900 demos [Verified] |
| Held to opportunity | 20 to 25% in strong programmes | Leadium, no sample [Secondhand] |
| SDR-qualified to opportunity | 58% | TOPO via Gradient Works [Secondhand] |
| SQL to close, outbound | 14%, derived not measured | Alba Talent [Secondhand] |
- Pull your own scheduler and CRM data: created timestamp, scheduled start, status, outcome, per claimed meeting. Use the scheduler, not the CRM, pick a period that ended two or more weeks ago, and strip duplicates, internal tests and spam.
- Reconcile booked against held. Count only meetings an external decision-maker joined. That gap is your first exposure.
- Dedupe by email, person and company across months. Watch for a contact counted twice, or a "new" meeting that is a reschedule of a no-show.
- Verify a real decision-maker at a target account, and demand call recordings. A REALQualified reviewer on G2 noted the agency "provided a recording of the call to verify the lead is poor quality" [Verified].
- Test domain and mailbox ownership: SPF, DKIM, DMARC, who registered the domains, who owns the mailboxes. On the agency's tenant you cannot keep them at exit, and the reputation dies with the contract.
- Demand the assets you paid for: sequences, lists, creative.
- Audit reply quality. Split positive replies from polite refusals and auto-responses. Saleshandy found 25 to 50% of replies were "interested" in strong campaigns; exclude out-of-office and tag negatives as referral, objection, not-now, unsubscribe [Secondhand].
| Failure mode | Signature | Evidence |
|---|---|---|
| Bookings that never hold | No held rate reported | Buyer via Prospeo: "$7,000/month... the agency had booked 14 meetings - and 9 of them no-showed" |
| Wrong person, wrong account | Title and account mismatch; low meeting-to-opportunity | CIENCE BBB, 23/10/2024: the buyer obtained "sworn affidavits from many of these so-called 'leads' affirming they had never heard of us" [Verified] |
| List broadened mid-engagement | Bookings climb as reply quality and ICP match fall | Inferred from the pattern above |
| Junior or offshore substitution | Unnamed reps on recordings; account-manager churn; sporadic activity from a "dedicated" person | CIENCE BBB, 26/08/2024: the rep "made calls for 3 days and then disappeared. We never ha[d] a consistent [rep] since signing." BBB, 14/10/2024: "we were assigned six different account managers." G2: "call center reps are all located offshore" [Verified] |
| Deliverability collapse hidden by volume | Sends rising, replies flat, spam complaints nearing 0.3% | Google and Yahoo thresholds [Verified] |
| Attribution capture | Agency-sourced meetings matching inbound form fills or existing relationships | Needs source stamping |
| Reporting-period gaming | End-of-month booking spikes, high share booked 8+ days out, then 23% no-show | Reply.io gradient [Verified] |
That fourth row is the complaint at the top of this page, dated and on the record three times over. You are not the first.
The standard pack is meetings booked, replies, sometimes a dashboard. Outbound System's Trustpilot listing sells "real-time dashboards where every metric is live" against rivals who "send PDF reports weekly or monthly," so the norm is a limited-metric periodic PDF [Secondhand]. Systematically absent: send versus delivered volume, bounce rate, spam-complaint rate, domain reputation, unsubscribe rate, list source and refresh cadence, per-rep attribution, held rate.
Deliverability is the one area governed by a real standard. Google and Yahoo bulk sender requirements, effective 1 February 2024 for senders of 5,000+ messages a day, mandate SPF, DKIM, DMARC, one-click unsubscribe (RFC 8058, enforced June 2024) and a spam-complaint rate below 0.3% calculated daily, Google's operational target being 0.1% [Verified]. Yahoo's Marcel Becker: "If you're a good sender, your spam rates will be well below 0.3%." Ask for that number by name.
The held quota splits by model: 16.0 introductory, 10.4 semi-qualified, 9.0 fully qualified. Stage 1 converted is a median of 6, down 43% since 2018, while pipeline per SDR rose to a $3.78M median from $2.83M in 2022, on bigger deals rather than more meetings [Verified].
Verdict, high confidence: 15 to 25 is aspiration. The "21 meetings a month, 62% conversion" line credited to Bridge Group sits at Gradient Works immediately beside a separate Operatix figure, which is how two vendors' numbers became one [Contradicted]. And 9 to 10 is a quota, not an achievement.
Bridge Group puts average SDR ramp at 3.0 months, the lowest since 2010 [Verified], and agency consensus is 60 to 90 days before data is reliable, first replies days after a two to three week warmup. Judging at week four is judging noise.
The circulating figures are not comparable. Walnut cites 38% appointment-to-opportunity for SaaS, no sample [Unsourced]. TOPO's 58% counts SDR-qualified leads. Leadium's 20 to 25% counts held meetings. Different denominators, stacked as if they were rungs of one ladder. Verdict, medium confidence: a disclosed-sample study fixing the "held" denominator does not exist [No evidence], so do not manufacture a single number and do not accept one.
That $134,000 is base $75K, OTE $25K, taxes and benefits $21K, tools $8.5K, onboarding $4.5K, and at 14.6 meetings a month derives $766 per booked meeting and $1,630 per qualified opportunity [Secondhand, derived]. Bridge Group confirms the pay magnitude: median OTE $80K [Verified]. The tech-spend figure of around $371 per SDR per month traces to an earlier edition of the same research, not the 2025 report [Secondhand].
Published figures omit tooling, data, domains and mailboxes, AE time to attend and qualify, and the cost of a no-show, which burns 15 to 60 minutes of AE capacity on top of the fee already paid. At 20% no-show, true cost per held meeting sits roughly 25% above cost per booked meeting. A quoted "$400 per meeting" understates the real figure by a wide margin.
No single no-show benchmark is correct, because the denominator decides the answer.
| Measurement | Figure | Sample |
|---|---|---|
| RevenueHero, 2 Dec 2024 | 6.5% (419 of 6,428) | One week, 15 industries; healthcare 0 of 179 [Verified] |
| RevenueHero, Aug 2025 | 13.5% median, 15.9% mean | 18 weeks, customers booking 50+ a month [Verified] |
| Reply.io, pre-2025 | 6.9% / 9.6% / 23.0% by lead time | 2,900 demos [Verified], possibly stale |
| "32% cold-booked 2025 vs 18% in 2020, Calendly" | Untraceable | The Calendly report located is 2023, n=1,241, different topic [Unsourced] |
| Ivris synthesis | Published range 6.5% to 28.1% | "Not a range you can average" [Secondhand] |
Verdict, high confidence that a single blended benchmark is wrong. For cold outbound with long booking lead times, expect 20%+. Ziellab restates the Reply.io shape, corroborating that gradient rather than the "32%."
Blended numbers mislead because everything segments. Belkins' 2025 analysis of 7.5M cold emails shows reply rate falling almost linearly with company size, 0.72% at 0 to 10 employees to 0.22% at 10,000+, and by seniority from founders 0.57% to C-level 0.42% to VPs 0.32% [Verified]. Cycle length moves the same way: 84 days median for B2B SaaS, 14 to 30 days under $15K ACV, 90 to 180+ above $100K [Secondhand].
Reply rates need the same scrutiny, and notice who publishes which. Belkins, an agency, reports 0.45% as replies over total sends. Instantly and Saleshandy, both tools, report 3.43% and 3 to 5% on looser bases [Verified]. Not a real conflict, just different denominators, the strictest coming from the agency. Any reply rate quoted without its denominator is unusable.
Four numbers likely to appear in your report should be struck: "15 meetings per SDR per month" [Unsourced], "21 meetings at 62% conversion" credited to Bridge Group [Contradicted], "no-show rose from 18% to 32%" credited to Calendly [Unsourced], and "multichannel drives 287% higher engagement" [Contradicted], which traces to Omnisend e-commerce purchase-rate data with the denominator silently switched. Five more, with the citation chain for each, are in the appendix.
Every "ideal reporting pack" I could locate is agency or tool content. A full-transparency standard from a neutral body does not exist [No evidence].
Numbers are only worth what the contract makes enforceable. The six terms to refuse, the two law firms' guidance on defining "lead" and assigning ownership of domains, lists and data, and the FTC context are in the evidence appendix.
The cycle outruns the evaluation window: 84 days median, 90 to 180+ for enterprise, and 74.6% of new-customer deals taking at least four months to close. Sixty to ninety days buys leading indicators only, so judge deliverability, positive-reply quality, held rate, meeting-to-opportunity and ICP match. To separate sourced pipeline from what would have closed anyway, hold 10 to 20% of the audience out of outreach. Unify: "without a control, 'would they have closed anyway' is unanswerable... 10 to 20 percent reserve is standard" [Secondhand]. No published incrementality methodology for agency outbound exists [No evidence], a category-wide weakness, mine included.
- What meeting-to-opportunity conversion rate is good?
- Use 20 to 25% of held meetings as a floor (Leadium). TOPO reported 58% of SDR-qualified leads become opportunities, but that is a different denominator. No disclosed-sample study fixes "held," so get the definition in writing first.
- How many qualified meetings per month is realistic?
- Roughly 9 to 10 held per full-time SDR, per The Bridge Group (6 February 2025, n=351 B2B companies), with the Stage 1 converted median at 6. The circulating "15 to 25 a month" carries no disclosed sample. Only 60% of reps hit quota in 2025.
- What KPIs should I ask an outbound agency to report?
- Booked versus held, ICP-match rate, meeting-to-opportunity, positive-reply rate, bounce and spam-complaint rate against the 0.3% cap, per-rep attribution, list source and refresh cadence, sourced versus influenced pipeline. What they omit tells you more than what they send.
- How do I calculate outbound cost per meeting?
- Retainer divided by held meetings, plus tooling, data, domains and internal AE attend-and-qualify time. Published cost per qualified B2B appointment is $550 to $1,700 (Clutch, 2025, via Leadium). Add roughly 20% no-show waste for true cost per held meeting.
- What benchmarks should I use to judge meeting quality?
- Show rate above 60% and meeting-to-opportunity of at least 20%. No-show ranges 1.2% to 18.1% by industry (RevenueHero, 6,428 meetings, Dec 2024) and 6.9% to 23% by booking lead time (Reply.io, 2,900 demos). Segment before comparing.
Run the seven-step audit. If held rate, ICP match and reply taxonomy come back clean, your agency is delivering and the suspicion was the cost of a category with no standards. If they come back thin, you have the reconciliation in writing, from your own systems, before renewal.
If you are still deciding whether an agency belongs in your stack at all, start with whether hiring a B2B lead generation agency is worth it.
Primary, sample disclosed
- The Bridge Group — 2025 SDR research report, 6 February 2025, n=351 B2B companies.
- RevenueHero — no-show analyses: December 2024 (6,428 meetings, one week, 15 industries) and August 2025 (18 weeks, customers booking 50+ meetings a month, no n disclosed).
- Reply.io — 2,900 booked demos, no-show by booking lead time.
- Belkins — cold email study, updated 26 June 2026, 7.5M emails across 2025 campaigns; appointment-setting pricing models.
- Instantly and Saleshandy — 2026 benchmark reports, 53.1M+ emails.
- CSO Insights — 2019 study; the primary release says "nearly 900 global sales leaders", the n=886 figure circulates via secondary reporting.
- Google and Yahoo — bulk sender requirements, effective 1 February 2024; RFC 8058 one-click unsubscribe.
- SiriusDecisions / Forrester — Demand Waterfall, debuted 2006 per Forrester, re-architected 2012; classic stage definitions per the 2015 SiriusDecisions eBook.
- Sprintlaw and Tech Attorneys — contract guidance on defining "lead" and asset ownership.
- BBB, G2 and Trustpilot — buyer complaint and review records, 2024–2025.
Secondary or commercial
- Alba Talent — fully loaded SDR cost derivation, 2026.
- Leadium — held-meeting conversion floors and the Clutch 2025 cost-per-appointment figure (the Clutch original is not locatable on clutch.co).
- Gradient Works — republishing Operatix and TOPO figures.
- Ivris Tech — no-show synthesis.
- Unify — incrementality and holdout guidance.
- Walnut — 38% appointment-to-opportunity claim.
- DemandNexus — pay-per-held-BANT-meeting billing.
- Prospeo, Outbound System, Ziellab — buyer stories, dashboards, corroboration.
- Landbase citing Omnisend — the "287% higher engagement" chain.
- MarketingCharts — reporting the CSO Insights cycle-length findings (74.6% of new-customer deals at four-plus months).
- ScaledMail and Instantly's help center — the 60–90-day and two-to-three-week-warmup figures.
- Also named: Optifai, Apollo, PhoneWagon, SalesAR, Clutch, Rocket Agents, Zeliq, Calendly, FTC.
All 24 with dates, URLs, samples and commercial flags, plus the nine-claim register, four unevidenced segments and eight gaps: evidence appendix.
Related reading
- Is Hiring a B2B Lead Generation Agency Worth It?The one independent data point is a LinkedIn poll: 7% of teams said outsourced SDRs really worked. I sell outbound, and I am publishing that anyway.
- Why Are All My Cold Emails Going to Spam?Four causes, four different fixes, and a twenty-minute free diagnostic to tell them apart. Worked from the provider docs and the RFCs, not from other blogs.
- Is GTM Dying? No — But One Tier of It Is Being GuttedGTM isn't dying — it's bifurcating. The commodity cold-outbound tier is being gutted while the strategy and systems tier grows. Which half are you in?