Quick answer
Forrester's 2025 State of Customer Obsession Survey found 74% of B2B and B2B2C organizations are already running AI agents and another 14% plan to, 88% combined. The same research ties failures to governance gaps, not the agents themselves. Ranked worst first, from what I see in outbound and GTM deployments: no ownership of what the agent can do, no audit trail on its actions, compliance added after launch, vendor benchmarks taken at face value, no human-in-the-loop checkpoint, and no spend ceiling. Fix the first two before you scale anything else.
The stat, and why adoption was never the hard part
Forrester published a blog post in March 2026 titled "The Future Of B2B GTM Isn't Human Versus AI", and the number that stuck with me was this one: 88% of B2B organizations are adopting or planning to adopt AI agents, per Forrester's State Of Customer Obsession Survey, 2025. Broken down, that is 74% already running agents in some form and another 14% planning to follow.
I read that number twice because it does not match the conversations I actually have with GTM leaders. Almost nobody tells me adoption is the problem. What they tell me about, when a deployment goes sideways, is a much narrower list of failures: an agent that sent something nobody approved, a CRM full of actions nobody can trace back to a decision, a compliance question that only came up after legal noticed. Forrester's own framing backs this up directly. The post argues that humans stay the differentiator precisely because governance gaps, not the technology, are what sink these projects.
So this is not a piece about whether to adopt AI agents in outbound and GTM. At 88%, that decision is largely made. This is a ranked list of the specific governance gaps I see cause the trouble, worst first, based on what actually breaks in outbound and GTM agent deployments rather than in the abstract.
How I ranked these six gaps
I ranked these by how often I see them cause a real incident, a bad send, a compliance exposure, a runaway bill, weighed against how expensive the gap is to close later versus early. A gap that is cheap to close before launch and expensive to unwind after go-live ranks higher than one that is merely annoying. This is not a scientific survey, it is a practical ordering based on deployments I have been close to and patterns I keep seeing repeated across the wider AI SDR and GTM agent conversation.
Tip. If you can only fix two things this quarter, fix #1 and #2. Everything below them is easier to retrofit once ownership and an audit trail exist.
#1: Nobody owns what the agent is allowed to do
Verdict: the gap that causes every other gap on this list. The most common failure I see is not a bad model output, it is that no single person or team can answer, in one sentence, what the agent is authorized to send, to whom, and under what conditions it should stop and ask a human first. When that ownership is fuzzy, every other control gets built inconsistently, because different teams assume different defaults.
The fix is boring on purpose: one named owner per agent, a written list of what it can and cannot do without a human sign-off, and a review cadence for that list. If you cannot name the owner in under five seconds when I ask, this gap is live in your stack right now.
#2: No audit trail on agent actions in the CRM
Verdict: the gap you only notice after something has already gone wrong. An agent that enriches a record, drafts a sequence, or updates a stage should leave a trail that says which agent, which prompt or trigger, and which human, if any, approved it. Without that, a bad send or a data quality problem becomes a forensic exercise instead of a two-minute lookup, and it makes it impossible to tell your own team, let alone a regulator, what actually happened.
This is also the gap that determines whether you can even measure the other five. An audit trail is the raw material for spotting patterns in the gaps below, not just a compliance checkbox.
#3: Compliance rules bolted on after launch
Verdict: cheap before launch, expensive after. This is the gap I see teams discover the hard way when a deadline lands on top of them. The EU AI Act's Article 50 disclosure duty and France's move to a mandatory opt-in rule for AI cold calling both take effect in August 2026, nine days apart, and teams that treated disclosure and consent as an outbound-copy afterthought are now retrofitting it into live sequences instead of a template built in from the start. I wrote a longer breakdown of that specific pair of deadlines separately, but the pattern generalizes: any jurisdiction-specific rule an agent needs to respect is far cheaper to encode as a hard rule before launch than to patch into thousands of already-running records after.
#4: Vendor benchmarks trusted without in-house checks
Verdict: common, and quietly expensive. Vendors report their own lift numbers, their own conversion parity claims, their own case studies. I do not think most of these are dishonest, but a number generated by the party selling the product is not the same as a number generated on your own list, your own ICP, your own compliance constraints. I have seen teams size a rollout, and a budget, off a vendor's own benchmark without ever running a controlled comparison against what they already had.
The fix is a short one: before you scale spend on any AI SDR or agent vendor's claim, run it against a holdout segment of your own pipeline first. It costs a few weeks. Skipping it costs a quarter of budget aimed at the wrong assumption.
#5: Autonomous-only, with no human checkpoint
Verdict: the gap the market has already priced in. The fully autonomous, remove-the-human model has consistently underperformed the hybrid model across the AI SDR data I have looked at this year, and Forrester's own framing agrees: the future is augmentation and orchestration guided by human intent, not agents running the full motion unsupervised. Autonomous-only setups tend to fail quietly, since there is no checkpoint positioned to notice degradation before it shows up in the pipeline numbers weeks later.
A single checkpoint, a human reviewing a sample of agent output before it scales, or approving anything touching a named-account or senior-buyer segment, catches most of what autonomous-only setups miss.
#6: No usage or spend ceiling on the agent tooling itself
Verdict: the smallest gap on this list, but the easiest to fix in an afternoon. This one is not about what the agent does to prospects, it is about what it costs you. Teams that skip usage alerts and hard ceilings on their own coding and orchestration agents find out about a runaway cost only when the invoice arrives. It ranks last because it is cheap and fast to fix, a spend alert and a hard cap, not because it is unimportant.
The six gaps, ranked
Here they are side by side, ranked worst first, with a rough sense of how the cost of fixing each one shifts depending on whether you address it before or after launch.
| Rank | Gap | Cost to fix early | Cost to fix late |
|---|---|---|---|
| 1 | No ownership of agent permissions | Low | Very high |
| 2 | No audit trail on agent actions | Low | High |
| 3 | Compliance bolted on after launch | Low | High |
| 4 | Vendor benchmarks untested in-house | Medium | Medium, plus wasted budget |
| 5 | No human-in-the-loop checkpoint | Low | Medium |
| 6 | No spend or usage ceiling | Very low | Low, but sudden |
Reading the table against your own stack
The pattern across all six rows is the same: every gap here is cheap before launch and expensive after. That is the actual lesson behind Forrester's 88% number. Adoption was never the hard part of this. Building the governance layer at the same time as the agent, instead of after the first incident forces the question, is.
Where the Forge stack fits into this
I run Salesforge for sending and Leadsforge for lead data, and the reason I have not run into most of these six gaps myself is that ownership and an audit trail were part of the setup from day one, not something I retrofitted. Every send has a traceable trigger, every list has a traceable source, and I know exactly which human approved what before it went out. None of that is unique to this stack, you can build the same discipline on any tooling, but it is the stack I check my own governance against, and it is why I default to it when a client asks what to run.
A short list for closing the top two gaps this quarter
If you only act on one thing from this roundup, act on this. Name one owner per agent this week, in writing, not verbally. Turn on logging for every agent-initiated CRM action if it is not already on, even if you have not decided what to do with the logs yet. Write, in one page, what the agent can do without a human, and what it must escalate. Review that page monthly, not annually, since agent capability changes faster than an annual review cycle can track. Everything else on this list gets easier once those three exist.
Key takeaways
- 88% of B2B orgs are adopting or planning to adopt AI agents, per Forrester's State Of Customer Obsession Survey, 2025, 74% already live and 14% planning to follow.
- Forrester's own framing ties project failures to governance gaps, not the agents themselves, which matches what I see in outbound and GTM deployments.
- The two highest-ranked gaps, unclear ownership and no audit trail, are also the cheapest to fix before launch and the most expensive to retrofit after.
- Vendor benchmarks are not a substitute for testing a claim against your own list and your own pipeline.
- A single human-in-the-loop checkpoint catches most of what fully autonomous deployments miss, and it is cheap relative to what it prevents.
FAQ
What percentage of B2B companies are adopting AI agents in 2026?
88%, per Forrester's State Of Customer Obsession Survey, 2025, cited in Forrester's March 2026 blog post "The Future Of B2B GTM Isn't Human Versus AI." That splits into 74% already adopting and 14% planning to.
Why do AI agent deployments in GTM and outbound fail if adoption is this high?
In my experience, and in Forrester's own framing, it is rarely the agent's capability that causes the failure. It is a governance gap: unclear ownership of what the agent can do, no audit trail on its actions, or compliance treated as an afterthought instead of a launch requirement.
What is the single highest-priority governance gap to close first?
Ownership. If no named person or team can state in one sentence what an agent is authorized to do and when it must escalate to a human, every other control in your stack gets built inconsistently around that ambiguity.
Should AI agents in outbound run fully autonomously or with a human in the loop?
Hybrid setups with a human checkpoint have consistently outperformed fully autonomous ones in the data I have reviewed this year. A single review point, sampling output or approving anything touching senior-buyer or named-account segments, catches most of what autonomous-only deployments miss.
How should a team verify a vendor's AI agent performance claims?
Run the claim against a holdout segment of your own pipeline before committing budget to it at scale. A benchmark generated by the vendor selling the product is not the same as a number generated on your own list under your own constraints.
Hlib Storchak · 2026-07-17 · ~10 min read