Quick answer
Outreach is pitching a data moat: 3 billion-plus signals and 33 million-plus weekly action-outcome pairings behind its "agentic AI" platform. I put that claim next to three other real 2026 vendor AI claims, Amplemarket Duo's self-run #1 ranking, 11x's Julian voice agent testimonial, and 11x's disputed Airtable and ZoomInfo customer logos, and ranked them by how much skepticism each deserves. None are proven fabrications. All four are self-reported. The useful move is not to dismiss them, it is to know which type of claim you are looking at before it justifies a price premium.
Outreach's pitch, and why I'm ranking it against three others
I'm Hlib Storchak. I build outbound systems for B2B founders and sales teams, and most of what follows comes from running this for clients who are the ones actually fielding these vendor pitches. I've booked 2000+ meetings for B2B clients, which mostly means I've sat in on more demos and renewal calls than I can count, and heard the same shape of claim wearing a different vendor's logo each time.
Outreach's current version of that claim is a data moat: more than 3 billion signals used to train its machine learning models, and more than 33 million action-outcome pairings captured every week across roughly 6,000 customers, all cited straight from Outreach's own product page. It is a real, specific, and reasonably impressive-sounding number. It is also exactly the kind of number that is impossible for a buyer to verify from the outside, which is what makes it worth a closer look, not a dismissal.
Rather than write one more single-vendor teardown, I picked three other AI claims making the rounds in 2026 and ranked all four side by side, from the one I'd trust the most to the one that should make you pause the longest before you let it move a price.
What Outreach is actually claiming
The specific numbers Outreach publishes are: 3 billion-plus signals used to train its models, 33 million-plus action-outcome pairings captured weekly, and roughly 6,000 customers generating that data. The pitch is straightforward: a bigger, longer-running dataset should make a better model, and a competitor starting today cannot replicate years of accumulated customer interactions overnight.
That logic is not wrong on its face. Data moats are real in machine learning, and scale genuinely matters for a lot of prediction problems. The gap between "this logic is sound in general" and "this specific number tells me the model will work for my team" is where the due diligence actually needs to happen, and it is the same gap that shows up in the other three claims below.
Not every vendor claim is the same species
Before ranking anything, it helps to separate what kind of claim you are actually looking at, because a "3 billion signals" claim and a "we ranked #1" claim fail in completely different ways. I sort vendor AI claims into three buckets: scale claims (we have more data or customers than you'd expect), performance claims (our thing converts at X%), and proof claims (named customers who use and endorse the product). Each one has its own way of going wrong, and its own way of being checked.
Scale claims are hard to verify and easy to inflate quietly, because the underlying dataset never leaves the vendor's servers. Performance claims are often true for the one account cited and misleading as a general expectation. Proof claims are the easiest to check and, when they turn out to be wrong, the most damaging, because a named customer can simply say "that's not true" in public, which is exactly what happens in the fourth claim on this list.
| Claim | What it says | Type | Independently verifiable? | My one-line verdict |
|---|---|---|---|---|
| Outreach | 3B+ training signals, 33M+ weekly action-outcome pairings across 6,000 customers | Scale claim | No, proprietary data | Plausible scale, but "signal" is undefined and pooled across every customer, not yours |
| 11x (Julian) | 50%+ demo-to-paid conversion, matched full-time rep conversion within 3 months | Performance claim | Partially, one published account | A real result for one customer, not a guarantee for yours |
| Amplemarket (Duo) | 219/231 points, #1 of 8 platforms, leads 9 of 10 categories in its own comparison | Performance claim | No, vendor built and scored its own test | Grading its own homework, useful for a feature checklist and nothing more |
| 11x (customer logos) | Airtable and ZoomInfo implied as customers on 11x's own marketing | Proof claim | Yes, and it did not hold up | Directly disputed by both named companies, the clearest warning of the four |
The four claims, at a glance
The table above is the full set I'm ranking. Two are self-run benchmarks or datasets a buyer cannot see inside. One is a genuine customer testimonial that a vendor is generalizing past its original scope. One is a named-customer claim that the named customers publicly denied. Ranking them side by side is more useful than judging any single one in isolation, because it shows you the shape of the problem, not just one instance of it.
Tip. A vendor that can tell you exactly what one "signal" or one "point" in its own scoring means is answering an engineering question. A vendor that repeats the number without defining it is answering a marketing question. Ask the definition question first, before you ask for the demo.
1. Outreach's 3 billion signals: plausible scale, unverifiable relevance
Outreach's number ranks first because it is the most internally consistent of the four. A platform running for 6,000 customers over multiple years plausibly does generate billions of interaction-level data points, and the 33 million weekly pairings figure is a reasonable-sounding derivative of that customer count. Nothing about the arithmetic looks fabricated.
What it does not tell you is whether the model trained on that pool actually improves your specific outcome, or whether "signal" means an email open, a reply, a booked meeting, a lost deal, or all of the above averaged together across every industry and deal size on the platform. A model trained on 3 billion generic B2B interactions is not automatically a model trained on 3 billion interactions like yours. That distinction is the whole ballgame, and Outreach's public materials do not resolve it.
2. 11x's Julian: a real result stretched into a general claim
11x's voice agent Julian is credited, on 11x's own product page, with helping one customer convert over 50% of demos to paid subscriptions and matching a full-time rep's conversion rate within three months. I have no reason to doubt that this happened for that specific account. It reads like a real customer testimonial, not an invented number.
The problem is scope, not honesty. A 50%-plus demo-to-paid number from one account, in one motion, at one price point, is not the same thing as "Julian converts at 50%-plus," full stop, which is how the claim tends to travel once it leaves the case study and lands in a sales deck. Ranking it second reflects that the underlying result is likely real, while the generalization built on top of it is doing more work than the evidence supports.
3. Amplemarket Duo's #1 ranking: grading its own homework
Amplemarket published a 231-point comparison of 8 AI sales platforms and scored its own Duo Copilot at 219 out of 231, first place, leading 9 of 10 categories. That is a genuinely detailed rubric, and building one at all is more rigorous than a lot of vendor content in this space.
It is still the vendor scoring itself against a rubric the vendor designed, with no disclosed independent audit of either the rubric or the scoring. A 231-point framework sounds objective because of the granularity, but granularity is not the same as neutrality. I rank it third because a self-graded exam, however detailed, tells you what the vendor thinks matters and how the vendor thinks it stacks up, which is useful market context and not evidence.
4. 11x's Airtable and ZoomInfo logos: the cautionary tale
This is the clearest case on the list, and the reason it ranks last. Airtable's logo appeared on 11x's own customer materials, and Airtable told reporters it was never actually a customer, having run a short test and passed. ZoomInfo separately ran a one-month pilot of 11x's product and, according to public reporting, concluded the product "performed significantly worse than our SDR employees" and did not proceed, then had its lawyers raise deceptive trade practice and trademark concerns over how the relationship was being represented.
A scale claim or a self-run benchmark can be argued about in good faith. A named customer publicly denying the relationship a vendor implied is not an argument, it is a fact check the vendor failed. If you are ever shown a logo wall as proof a category works, this is the exact scenario it is meant to protect against, and the exact scenario it failed to prevent here.
Key takeaways
- Outreach's 3B-plus signal claim, Amplemarket's self-run #1 ranking, 11x's Julian testimonial, and 11x's disputed customer logos are all self-reported, and none of that makes them automatically false.
- Scale claims, performance claims, and proof claims fail in different ways. Sort a vendor's pitch into one of the three before you decide how to check it.
- A named-customer claim is the easiest type to verify and the most damaging when it fails, which is exactly what happened with 11x's Airtable and ZoomInfo logos.
- A self-run benchmark, however detailed the rubric, is the vendor grading its own homework. Treat the score as a feature inventory, not independent proof.
- Before any of these claims justifies paying more, run the actual cost math against your own numbers, not the vendor's headline figure.
The pattern underneath all four
Every one of these claims originates from the vendor, about the vendor, published by the vendor. That is not a scandal, it is simply how vendor marketing works, and it applies to every company selling into this space, including ones I recommend to clients. The mistake is not that vendors publish self-favorable numbers. The mistake is a buyer treating a self-reported number as due diligence that has already been done, rather than as the starting point for due diligence that still needs to happen.
The other pattern worth noticing: the claim that failed publicly, the customer logos, was also the easiest one to check. Scale claims and self-run benchmarks are hard to verify because the data lives inside the vendor. Named-customer claims are easy to verify because the named customer can simply be asked. If a vendor's pitch leans hard on logos or named case studies, that is actually the good news, because it is the one part of the pitch you can fact-check directly instead of taking on faith.
The price-premium test: does the claim change your math
None of this matters much if the "agentic" or data-moat feature is free. It matters a great deal when it comes with a renewal price increase, which is the position a lot of buyers are in right now. Here is the assumption-based way I run that math with clients, using placeholder figures you should replace with your own before deciding anything.
Assumptions (swap in your real numbers): a current per-seat cost of roughly €70 to €120 a month, a vendor asking for a 25% to 40% premium to unlock the "agentic" tier, and a 10-seat contract. The formula is: extra monthly cost = seats × base cost × premium percentage.
| Input (your assumption) | Example range |
|---|---|
| Current cost per seat, per month | €70 to €120 (assumption, replace with your invoice) |
| Claimed "agentic" premium | 25% to 40% (assumption, replace with the quoted uplift) |
| Seats on the contract | 10 (assumption, replace with your seat count) |
| Extra monthly cost (calculated) | roughly €175 to €480 |
| Assumed value of one incremental meeting | €150 to €400 (assumption, replace with your pipeline math) |
| Extra meetings per month needed to break even (calculated) | roughly 0.5 to 3 |
With those placeholder ranges, breaking even needs somewhere between half an extra meeting and three extra meetings a month attributable specifically to the new claim, not to the platform overall. That is a low bar for some teams and a real bar for others, which is the point: the number that should move your decision is your own break-even meeting count, not the vendor's headline signal count or ranking score.
Six questions before a claim earns a premium
I run the same short list on any vendor claim before I let it justify a renewal increase for a client, regardless of which of the four categories above it falls into.
- What exactly does the unit in the claim mean, a "signal," a "point," a "conversion," defined precisely enough that I could reproduce the count myself?
- Is the claim about a pooled average across every customer, or about accounts that look like mine in industry, deal size, and motion?
- If it is a self-run benchmark, who designed the rubric, and is there any published methodology I can check against a competitor's own version of the same test?
- If it is a named-customer proof point, can I contact that customer directly, or has anyone already tried and reported back?
- What is the actual extra cost, in my currency, on my seat count, not the percentage the vendor leads with?
- How many extra meetings, opportunities, or hours saved does that extra cost require to break even, using my numbers, not a case study's?
A vendor with a real claim answers all six without flinching. A vendor mid-pitch sometimes cannot, not because the claim is false, but because the marketing usually outruns the documentation that would let a buyer verify it.
What I actually check when a client wants to buy the story
When a client tells me they are ready to sign because a vendor's AI pitch impressed them, the first thing I do is not argue with the pitch. I ask them to run the price-premium math above with their own contract numbers, and I ask the vendor, on the client's behalf, to define the unit behind whatever headline number sold the room. More often than I would have expected, that single question, what exactly counts as one of these, slows the deal down enough for the client to negotiate the premium down or ask for a longer pilot before committing to it.
That is not skepticism for its own sake. Data moats, self-run benchmarks, and customer testimonials can all be genuinely true and still not be the reason to pay more, if the math on your own contract does not clear the bar. Getting that math in front of the client before the renewal call, not after, is the part of this that actually changes outcomes.
My honest take
I do not think any of the four vendors here are acting in bad faith, including 11x, whose customer-logo situation is the one that looks worst on paper. Every serious platform in this category is under pressure to prove an AI story fast, and self-reported numbers are the quickest story any of them can tell. What I push back on is a buyer treating the number itself, 3 billion, 219 out of 231, 50%, as the finish line of due diligence rather than its starting point.
The ranking above is not a verdict that Outreach, Amplemarket, or 11x's Julian are lying. It is a reminder that four different flavors of self-reported claim carry four different amounts of risk, and the cheapest way to tell them apart is to ask what the unit means and who else can confirm it, before the number is allowed anywhere near your renewal price.
FAQ
Is Outreach's 3-billion-signal claim fabricated?
I have no evidence it is fabricated. It comes from Outreach's own product materials, which makes it a real number reported by an interested party, not an independently audited figure. That distinction, not the size of the number, is what should shape how much weight you give it.
How do I know if a vendor's self-run benchmark, like Amplemarket's 219-out-of-231 ranking, is trustworthy?
Ask who designed the rubric and whether the methodology is published in enough detail that a competitor could run the same test on itself. A detailed scoring system is still the vendor grading its own homework unless an independent party ran it.
Should a single customer testimonial like 11x's 50% demo-to-paid number change my buying decision?
Treat it as evidence the result is possible, not evidence it is typical. Ask for the account's industry, deal size, and motion, and compare that against your own before assuming the same number applies to you.
What does the Airtable and ZoomInfo dispute with 11x actually prove?
It proves that named-customer claims are the easiest type of vendor claim to check and the most damaging when they fail. Both companies publicly disputed the relationship 11x's marketing implied, which is the clearest of the four warning signs in this article.
How do I decide if an "agentic AI" price premium is actually worth paying?
Run the break-even math on your own contract: extra monthly cost divided by the value you assign to one incremental meeting or hour saved. If the vendor's claim would need to deliver an unrealistic number of extra wins a month to clear that bar, the premium is not worth it yet, regardless of how the claim is worded.
Hlib Storchak · 2026-07-24 · ~10 min read