Quick answer
Not automatically. A 10,000-email study published by Prospectory in March 2026 found AI-written cold emails replying at 8.2% against 11.7% for human-written ones, with human emails also booking more meetings and closing more deals, per Prospectory's study. The lift vendors advertise shows up only when AI is paired with real signal and a tight list, not from swapping a human writer for a model on the same generic send.
The claim vendors make
Open a demo for almost any AI SDR or cold email tool right now and you will hear a version of the same pitch: AI writes a more relevant email than a human ever could, because it reads more of a prospect's context before it drafts a line, and that relevance shows up as a higher reply rate. It is a reasonable-sounding claim. It is also exactly the kind of claim I get asked to verify before a client signs a contract, because "AI personalization lifts replies" is doing a lot of work in a sentence that usually has no attached source.
Why this matters. If the underlying claim is wrong, you are not just wasting a subscription. You are training a team to trust vendor benchmarks over their own send data, which is a habit that gets expensive the next time a claim really does not hold up.
A 10,000-email study that found the opposite
The most specific dataset I could find on this exact question is not from a tool vendor selling AI writing. It is a study Prospectory published on March 9, 2026, testing 10,000 cold emails, 5,000 generated by GPT-4 and 5,000 written by human SDRs with at least two years of experience, sent over three months. Per Prospectory's own writeup, AI-written emails replied at 8.2% against 11.7% for the human-written set, a 43% gap in favor of the humans. That is the reverse of what most personalization pitches promise, and it is worth sitting with rather than skipping past.
The numbers, side by side
Reply rate alone does not tell the full story, so here is the same study broken out by the metrics that actually determine whether a reply turns into revenue.
| Metric | AI-written | Human-written |
|---|---|---|
| Overall reply rate | 8.2% | 11.7% |
| Substantive reply (not a one-liner) | 43% | 62% |
| Meeting booked rate | 1.9% | 3.4% |
| Close rate on booked meetings | 12% | 19% |
| Average deal value | $31,000 | $47,000 |
| Cost per email sent | $0.12 | $4.80 |
The cost line is why this is not a clean "humans win" story either. AI email cost roughly 2.5% of what a human-written email cost in the same study, per Prospectory's cost breakdown. A tool that replies worse per email but costs forty times less per email is a different decision than a tool that simply underperforms, and the right call depends on what you are actually optimizing for on a given list.
Why the gap shows up where it does
The industry breakdown in the same study is the most useful part. In SaaS, AI and human performance were close, 9.8% against 10.2%, a gap small enough to call a wash. In healthcare, human-written emails replied at 14.8% against 5.4% for AI, and in financial services humans replied at 13.2% against 6.9% for AI. Those are the two most regulated, most relationship-driven verticals in the set, and they are exactly where a generic model's draft reads as generic to someone who can tell. The gap is not really "AI versus human." It is "AI writing without deep account context versus a human who already knows the account," and that difference in inputs, not the writer, is what the reply rate is actually measuring.
This is not the first claim that did not survive contact with data
I have gone through this exercise with clients before on other AI SDR claims, autonomous-agent ROI numbers, unit economics slides, churn figures, and the pattern repeats: a headline number from a vendor's own case study looks strong until you ask what it is being compared against, and the comparison usually flatters the tool being sold. That does not make every vendor number false. It means a number with no named methodology, no sample size, and no independent replication is a marketing input, not a fact, until you have checked it against something else.
So is AI personalization actually useless?
No, and the same study does not say that either. It says AI writing without strong signal underperforms a good human writer on a cold list. That is a narrower claim than "AI personalization does not work," and it lines up with what most operators actually see: an AI draft that pulls only a name, a company, and a generic pain point reads like every other AI draft a prospect has already ignored twice that week. An AI draft that is fed a real trigger, a funding round, a job change, a specific product signal, reads differently, because the input changed, not because the AI got smarter overnight.
What actually moves the number: signal, list size, volume
Across the benchmark data I could find, the consistent driver of reply rate is not who or what wrote the email. It is targeting depth. Instantly's 2026 benchmark report, drawn from billions of aggregated sends, puts the platform-wide average reply rate at 3.43%, with top-quartile senders at 5.5% and elite campaigns clearing 10%, per Instantly's Cold Email Benchmark Report. The senders clearing 10% are not doing it with a better model. They are doing it with a smaller, tighter list matched to a real signal, which any writer, human or AI, can turn into a better email. Volume and writer quality both matter less than the size and precision of the list you are writing to.
Five questions to ask before you believe a personalization claim
When a vendor, a case study, or a LinkedIn post tells you AI personalization lifted replies by some percentage, run it through these before you act on it.
- Compared against what? A lift over no personalization at all is a different claim than a lift over a skilled human writer.
- What was the sample size and time window? A few hundred sends over two weeks is noise. Thousands over a full quarter is a real signal.
- Whose data is it? A vendor citing its own customers' results has an obvious incentive. Independent or academic data carries more weight.
- Does it separate signal quality from writer quality? Most lift comes from better targeting. A claim that does not isolate this is probably crediting the AI for the list.
- Does it report downstream metrics, not just replies? A reply is not a meeting, and a meeting is not a closed deal. Ask for all three before you believe any one of them.
How I verify a claim like this on my own list before I trust it
I do not take any vendor's benchmark as the number I plan around, mine included. Before I let a claim change how I run a client's outbound, I split a real segment of their list in half, run the same offer and the same targeting through both an AI-drafted and a human-drafted version, and let it run for at least two to three weeks before reading the result. That window matters because reply rates on a fresh sequence swing a lot in the first few days and settle later. The number I trust is the one from that split test on that specific list, not the number from someone else's case study, because list quality, industry, and offer all move the result more than the writer does.
Where the Forge stack fits in this decision
This is also why I keep the sending and the writing decisions separate in my own stack. I run Salesforge for sending and its AI drafting sits on top of infrastructure and deliverability tooling that is maintained by people whose job is exactly that, so a bad AI draft does not also risk burning a domain. Leadsforge is where the signal comes from, the enrichment and intent data that determines whether an AI draft has anything real to personalize against in the first place. The lesson from the Prospectory study is not "avoid AI drafting." It is "do not expect an AI draft to outperform a good human one on a thin, generic list," and a stack that keeps sending, data, and drafting cleanly separated is what makes it possible to test that honestly instead of guessing.
Key takeaways
- A 10,000-email 2026 study found AI-written cold emails replying at 8.2% against 11.7% for human-written ones, with human emails also booking more meetings and closing more deals, per Prospectory's published study.
- AI email in that study cost roughly 2.5% of what a human-written email cost, so the tradeoff is not simply "worse," it is cheaper and worse on a cold, low-signal list.
- The gap was smallest in SaaS and largest in relationship-driven, regulated verticals like healthcare and financial services, where generic AI drafts read as generic fastest.
- Across benchmark data, targeting depth and list size predict reply rate far more reliably than who or what wrote the email.
- Before trusting any personalization claim, check what it was compared against, the sample size, whose data it is, whether it isolates signal from writer quality, and whether it reports downstream metrics beyond replies.
- The only number worth planning around is one from a split test on your own list, run for at least two to three weeks.
My take
I do not think this study proves AI personalization is a bad bet. I think it proves the specific claim "AI personalization lifts reply rates" is not automatically true, and treating it as automatically true is how teams end up disappointed by a tool that was never actually tested against their own list. The operators I see getting a real lift from AI drafting are feeding it real signal and checking the result against a human-written control, not trusting the number on a landing page. That is a cheap habit to build and an expensive one to skip.
FAQ
Does AI-written cold email really reply worse than human-written email?
In a 10,000-email study published by Prospectory in March 2026, yes: 8.2% for AI-written versus 11.7% for human-written, with human emails also booking more meetings and closing more deals. That is one study on one dataset, not a universal law, but it is the most specific head-to-head data I found on this exact question.
Is AI personalization ever worth using then?
Yes, when it is fed real signal, a funding event, a job change, a specific product trigger, rather than just a name and a company. The study shows AI underperforming on generic, low-signal drafting, not that AI drafting can never work.
What actually drives a higher reply rate if not the writer?
Targeting depth and list size, based on benchmark data from Instantly's 2026 report showing top-quartile senders at 5.5% and elite campaigns above 10% against a 3.43% platform average. The senders at the top are working smaller, tighter, signal-matched lists, regardless of who wrote the email.
How do I check a vendor's personalization claim before I believe it?
Ask what it was compared against, the sample size and time window, whose data it is, whether it isolates signal quality from writer quality, and whether it reports meetings and closed deals, not just replies. Then run your own split test on a real segment of your list for at least two to three weeks before trusting the result.
Should I stop using AI to draft cold emails?
No. Stop trusting a vendor's reply-rate claim without checking it against your own list first. AI drafting paired with strong signal and a tight list is a different tool than AI drafting on a cold, generic list, and the studies available so far are mostly testing the second case.
Hlib Storchak · 2026-07-13 · ~10 min read