AI lead generation gets sold as outbound at scale, and outbound is the half that broke. Two constraints, both documented, neither fixable by a better model. First, research is metered and sending is not: Clay's own price list runs $446 a month for 40,000 actions and 6,000 data credits, data credits start at $0.05 each, overage carries a 30 percent premium, and Clay moved token-heavy models to variable pricing because deep research stopped fitting a flat rate. Woodpecker's platform data across more than 20 million emails shows the trap directly, with advanced personalization replying at 17 to 18 percent against 7 to 9 percent for basic, while reply rate falls from 5.8 percent under 50 contacts to 2.1 percent above 500. Second, Google has gated senders above 5,000 messages a day since February 1, 2024 and Microsoft since May 5, 2025, and published reply averages now sit at 3.43 percent. None of that applies to a person who already wrote to you. The inbound half is where agent work clears, and the honest speed evidence is one study of 15,000 leads, not the recycled conversion table this post used to carry.
AI lead generation is almost always sold as one thing: outbound at scale, with a model writing the emails. That half broke, and not because the models are bad. It broke on two constraints a better model does not touch, both documented in public by the vendors involved. The half that still works is the one nobody demos, and it lives further down the sales funnel than the top of it.
The consensus advice is to write better prompts, buy better intent data, and send more. That advice assumes the bottleneck is message quality. Checked against what Clay charges, what Woodpecker measures, and what Google and Microsoft enforce, the bottleneck is economic and administrative. Research costs money per prospect and does not get cheaper with volume. Delivery into a stranger's inbox is rate limited by two companies. Neither of those is a prompt problem.
The short version
- Research is metered. Sending is not. That asymmetry decides what gets automated, and it points automation at the wrong half.
- Clay's Growth plan is $446 a month for 40,000 actions and 6,000 data credits. Data credits start at $0.05 each and overage carries a 30 percent premium.
- Clay moved token-intensive models to variable pricing based on real token consumption. That is a vendor confirming deep research does not fit a flat rate.
- Woodpecker's data: advanced personalization replies at 17 to 18 percent, basic or none at 7 to 9 percent.
- Same dataset: reply rate falls from 5.8 percent under 50 contacts to 2.1 percent above 500. The channel gets worse exactly where you add volume.
- Google gated senders above 5,000 messages a day on February 1, 2024. Microsoft followed on May 5, 2025 at the same threshold.
- Published reply averages now sit at 3.43 percent across two independent platform datasets.
- None of the above applies to a reply in an open thread. An inbound lead has already replied, which is the part outbound spends all its money trying to buy.
- So the agent opportunity is inbound follow-up and qualification, where the research is free, the gate does not exist, and the scarce input is response time.
Why the outbound half of AI lead generation underdelivered
Two constraints, both documented, neither fixable by a better model. The research that makes cold outreach work is metered per prospect and does not get cheaper as volume rises. The delivery of that outreach is rate limited by Google and Microsoft. Automating volume automates the half that got cheaper, which is also the half that stopped working.
Research is metered. Sending is not.
Woodpecker's platform data states the trap without needing an argument on top of it. Across more than 20 million cold emails from over 1,000 customers in 52 countries, updated June 23, 2026, advanced personalization runs a 17 to 18 percent reply rate against 7 to 9 percent for basic or none. In the same dataset, over the same period, on the same platform, reply rate falls as list size rises: 5.8 percent under 50 contacts, roughly 3 percent between 200 and 500, and 2.1 percent above 500. Two findings from one source, and they point in opposite directions.
Read together they describe an inverse relationship rather than a tooling gap. The input that lifts reply rate the most is the input that scales the worst, because personalization at depth is per prospect work that does not get cheaper per prospect when you do more of it. Volume, meanwhile, actively degrades the outcome. A channel where the best lever resists scaling and scaling makes the result worse is not behaving like a channel you should automate for throughput.
The market has already priced this, which is the part worth checking yourself. Clay's published pricing is the closest thing to a list rate for research at scale: Launch at $167 a month for 15,000 actions and 3,000 data credits, Growth at $446 a month for 40,000 actions and 6,000 data credits. Data credits are the ones spent on real research, described on Clay's own page as what you use when buying data across its 150-plus partners. They start at $0.05 each. Actions start under a cent. Buying beyond your monthly allocation on Launch or Growth carries a 30 percent premium.
The most telling line on that page is not a price. Clay states that token-intensive models now carry variable pricing based on actual token consumption, while about 80 percent of models still cost a flat number of data credits per task. A vendor does not move a resource from flat to metered for fun. It does that when the flat rate stopped covering what the resource actually costs, which is a vendor saying out loud that letting a model do genuinely deep research on a prospect is expensive and gets more expensive the deeper it goes.
One caveat on reading those two cuts together: they are separate slices of the same dataset and their absolute levels do not reconcile with each other, since the 7 to 9 percent basic-personalization figure is not measured on the same campaigns as the 2.1 percent large-list figure. Woodpecker does not publish the cross-tabulation, so nobody can say what a thin-research campaign of 800 contacts replies at. What both cuts agree on is the sign: depth up, reply up. Volume up, reply down.
So the arithmetic every team hits looks like this. A send costs effectively nothing. A well-researched send costs credits and tokens that scale linearly with prospect count and refuse to get cheaper per prospect at depth. You can research properly and reach few people, which surrenders the throughput that was the entire reason to automate. Or you can reach many people on thin research, which moves you down on both of Woodpecker's axes at once. Most tools take the second fork, because it is the one the pricing pushes them toward. Anyone comparing AI SDR tools for demand generation should ask which fork a given product took before reading its feature list, and how a GTM platform and Clay differ on this is largely a question of whether the research is bought per prospect or already sitting in your own data.
The inbox belongs to Google and Microsoft
The second constraint is not economic. It is administrative, and it is documented on Google's own support pages rather than inferred from a decline curve. Google and Yahoo announced joint bulk sender requirements in October 2023, effective February 1, 2024. Google's requirements page states that senders of more than 5,000 messages a day to Gmail accounts must set up SPF and DKIM plus DMARC, must support one-click unsubscribe on marketing mail, and must keep spam rates in Postmaster Tools below 0.30 percent, with Google recommending below 0.10 percent. Microsoft imposed its own authentication requirements at the same 5,000 messages a day threshold on May 5, 2025, and non-compliant mail is rejected outright rather than filed in a junk folder.
Published reply averages moved in the direction you would expect. Instantly's Cold Email Benchmark Report 2026, updated January 12, 2026 on data from January 1 to December 18, 2025 across what it describes as billions of interactions, puts the overall average at 3.43 percent. Woodpecker independently reports 3.43 percent, down from 5.1 percent in 2024. Those two landing on the same figure to two decimal places is a coincidence rather than a second measurement confirming the first is precise, and the honest version of that whole benchmark argument is worked through in what signal-based selling got right and wrong.
The precise level matters less than who sets the ceiling. Both gates are keyed to volume, at 5,000 messages a day, and enforced on authentication and complaint rate. A more relevant message helps the complaint rate, which is real and worth doing. It does nothing about the existence of the gate. Any lead generation strategy whose payoff scales with send volume has its ceiling administered by two companies that have already lowered it once and gave about four months of notice when they did.
What both constraints have in common
Both are about the send. The metering is on manufacturing a reason for a stranger to reply, and the gate is on delivering it. Every dollar and every rate limit in outbound sits on one job: buying a reply from someone who did not ask. That framing is what makes the next part obvious, because there is a category of lead where the reply already happened and neither constraint applies.
The inbound half is where agent lead generation clears
On an inbound lead both constraints vanish. The research is already done, because the prospect did it and told you what they wanted by requesting a demo, reading the pricing page twice, or comparing plans. And nobody filters a reply to a person who wrote to you first, because it is a reply in a thread they opened. What is scarce on that side is response time, which is the one input an agent holds at no marginal cost.
Stated as an asymmetry, it is stark. Outbound spends metered research credits and burns sender reputation to manufacture a reply that arrives 3.43 percent of the time. Inbound starts with the reply in hand. A demo request is a reply that already happened. The only remaining questions are whether anyone answers it, how fast, and whether the answer reflects what that specific person actually looked at. Those are three questions about follow-through, not three questions about reach, and they are the reason inbound and account-based motions end up wanting different tooling than a sequencer.
| What the work needs | Cold outbound | Inbound follow-up |
|---|---|---|
| Research on this specific prospect | Bought per prospect. Data credits start at $0.05 each, overage at a 30 percent premium | Already collected as first-party behavior. Free at the point of capture |
| Permission to reach the inbox | Gated above 5,000 messages a day by Google since February 1, 2024 and Microsoft since May 5, 2025 | Not gated. It is a reply in a thread the prospect opened |
| Reply rate at the start of the work | 3.43 percent average across two platform datasets, 2.1 percent on lists above 500 | The reply is the trigger. It has already happened |
| Effect of adding volume | Degrades it. 5.8 percent under 50 contacts falls to 2.1 percent above 500 | No effect. Each lead is handled once, and volume is set by demand, not by sending |
| The scarce input | Money and sender reputation | Minutes, plus knowing what the person actually did on your site |
| Who else has the same input | Everyone buying the same third-party intent feed | Nobody. It is your own first-party data |
That last row is the quiet one. Third-party intent data is sold simultaneously to every competitor in your category, which is why a surging account gets five near-identical emails in the same week. First-party behavior on your own site and product is not for sale to anyone, and it is more specific than any intent feed can be, because it records what the person did rather than what a cooperative inferred about their company. The difference between first-party website data and scraped intent is the same argument applied to the input side.

What the speed data actually supports, and what it does not
The direction is well established. The precise numbers circulating are not, and this post used to carry one of the bad ones. The single most defensible source is the Lead Response Management study, run with Professor Oldroyd at MIT across more than 15,000 leads and more than 100,000 call attempts at six companies over three years of data.
Its central finding, published on the study's own site, is a 21-fold decrease in the odds of qualifying a prospect when response time stretches from 5 minutes to 30 minutes, with a fourfold drop occurring between 5 and 10 minutes alone. Across the first hour, contact odds fall more than tenfold and qualification odds fall more than sixfold. That is a decay curve with a real sample behind it, and it is steepest in the window a human queue cannot cover.
One honest caveat: this data is close to two decades old, from the same original Oldroyd research that later informed the 2011 HBR piece above. Nobody has re-run a study this rigorous on speed-to-lead since, which says something about how settled the field considers the direction, but a reader should know the underlying numbers predate the current AI-tooling era rather than assume they're fresh.
Now the correction. An earlier version of this post reproduced a five-row table putting lead-to-opportunity conversion at 21 percent under 5 minutes and 2.3 percent at 24 hours or more, attributed to an aggregator blog. That page, opened directly in July 2026, does not contain the table. It carries comparative pairs sourced to the Oldroyd study, to Velocify, and to a paywalled 2011 Harvard Business Review article, and other aggregators quoting the same HBR piece give the average first-response figure as 42 hours in some posts and 47 in others. The table has been removed rather than re-sourced, because a five-row conversion matrix nobody can trace to a primary study is not evidence, however well it formats.
What survives is enough to act on. Qualification odds decay fastest in the first 10 minutes, and a 5-minute response window is not a staffing problem a person solves. It is nights, weekends, the 40 minutes your one rep is on another call, and the Tuesday when four demo requests land at once. Rules-based lead routing already tried this and only solved delivery: it routes a record, it does not read what the visitor looked at, and it does not draft anything. The gap between routing a lead in 30 seconds and answering it well in five minutes is exactly the shape of work an agent does.
Where Intempt sits, and what one workspace changes
The honest claim is structural rather than magical. If the website behavior, the product usage, and the deal record live on one profile in one system, then the research step on an inbound lead is a lookup instead of a purchase, and the follow-up can reference what the person actually did rather than what their company allegedly cares about. That is why the agentic GTM platform framing is a data-model claim before it is a feature claim. The SDR reads that profile, drafts the follow-up, and stages it. You review it and send it, which is the part that keeps the reply sounding like it came from a person, because it did.
What it does not change: nothing here manufactures inbound volume. A company with 20 visitors a month has almost nothing to follow up on, and for that company outbound is the only option available, constraints and all. This argument is about where to point automation once inbound exists, not a claim that outbound is dead. The honest map of where agent work pays off puts numbers on which parts of the funnel actually move, and what fully autonomous AI SDRs get wrong covers the autonomy question separately from the inbound-versus-outbound one.
The named intent-data market, and how much of it you need
Signal-based prioritization has a real, named vendor set, and it is worth knowing before deciding how much of it to buy. Forrester's Wave for B2B intent data, Q1 2025 names six leaders: Intentsify, 6sense, Bombora, Informa TechTarget, ZoomInfo, and Demandbase. 6sense de-anonymizes account-level traffic and maps it against third-party signals to predict buying stage. Bombora runs a publisher cooperative that flags accounts spiking above baseline on a topic. Clearbit was acquired by HubSpot in 2023 and now runs inside HubSpot as Breeze Intelligence.
The practical read on that list follows from the two constraints above. These providers are good at telling you an account is in market, and the act that finding points you toward is usually a cold email into a gated inbox that every competitor on the same feed is also sending. Stacking signals helps: a topic surge plus a review-site comparison view plus a first-party pricing-page visit is more reliable than any single provider's score. But the first-party visit is the one that carries an unbroken thread back to a specific person, and it is the one you already own.
The AI Marketing Campaign Generator turns a goal, audience, and budget into a campaign plan and timeline in seconds, which is a useful first pass at the demand-generation side that has to exist before there is any inbound to follow up on. It runs on a free Intempt plan.
The limitation that survives all of this: data quality
- Scoring is only as good as the data under it. Incomplete CRM fields degrade ranking quality directly.
- Inconsistent UTM tagging means the model ranks on signals that are missing or mislabeled for a real share of traffic.
- A confidently wrong ranking is worse than none, because a sales team will trust it and work it.
- Fix tagging and required fields before layering scoring on top, not after the first bad quarter.
- This applies harder on the inbound side, because a pricing-page visit you failed to tag is a signal you cannot act on at all.
What to actually do
- Count last month's inbound demo requests, trials, and pricing-page sessions. If that number is above zero, this is where the cheapest lift is.
- Measure your median and worst-case first response time on those leads. Measure it from the prospect's action, not from when the record appeared in the CRM.
- Audit UTM tagging and required CRM fields. Unglamorous, and it is the real ceiling on any scoring you add later.
- Pick three to five first-party intent signals rather than scoring on everything. Pricing-page visits, repeat sessions, feature usage, and demo requests are enough to start.
- Set a response target for high-signal inbound and staff the gap with an agent that drafts, not one that sends. The decay curve is steepest in the first 10 minutes.
- Only then decide how much third-party intent data to buy, and price it against the gated inbox it points you at.
What this post does not claim
- Not that outbound is dead. It works at low volume with deep research, which is exactly the configuration you cannot automate for throughput.
- Not that Clay is overpriced. Its pricing is the clearest published statement of what research at scale costs, which is why it is quoted here rather than criticized.
- Not that Google and Microsoft are wrong to gate bulk senders. The rules cut real abuse. They are only a problem for a strategy whose payoff scales with volume.
- Not that the 3.43 percent figure is precise. Two large platform datasets agree on it and on the direction. Neither establishes the cause.
- Not that intent-data providers do not work. They identify in-market accounts accurately. The question is what act their finding points you toward.
- Not that an agent should send on its own. Every claim here assumes drafts staged for a human, which is also what keeps the reply sounding human.
- Not that any of this manufactures demand. Inbound follow-up cannot beat outbound at a company with no inbound.
Read the primary sources rather than the roundups, this one included. Clay publishes its rates. Woodpecker publishes its reply rates by personalization depth and list size. Google publishes the sender requirements and the complaint ceiling. The Lead Response Management study publishes its sample. Between them those four pages describe the whole problem without anyone needing to characterize a competitor, and they point the same direction. The half of the funnel everyone automated is the half that is metered on the way in and gated on the way out. The half worth automating is the one where the prospect already replied, which is why AI lead generation pays off fastest on inbound follow-up rather than on reach.
Frequently asked questions. Answered.
That is how it gets sold, and it is the half that stopped working. Two constraints sit on automated cold outreach. Research at depth is metered per prospect and does not get cheaper as volume rises, so scale forces you into shallow research. And delivery into a stranger's inbox is gated by Google above 5,000 messages a day since February 1, 2024 and by Microsoft since May 5, 2025. Neither constraint applies to following up on someone who already raised a hand.






