← All articles
CX Metrics & ROI

Cost Per Outcome, One Year Later: What the Numbers Look Like at Scale

A year ago we argued you should measure your AI front line by cost per booked outcome, not deflection. Here is what held up once real deployments ran at scale — and the honest range operators should plan against.

The Verbose CX teamJuly 26, 2026 · 8 min read

A year ago we made a single argument: stop scoring your AI front line on how many contacts it kept away from a human, and start scoring it on what it actually produced — the cost per booked outcome. It was easy to write when the deployments were young. The interesting question was always what the number would look like after twelve months of real traffic, escalations, and the slow grind of keeping an agent good. This is that follow-up, with the parts that held up and the parts that didn’t.

One argument again, because the framework only earns its keep if it survives the boring middle of a deployment: the cost-per-outcome number is real, it is defensible to a CFO, and it lands in a narrower and less flattering range than the launch decks promised. If you are past the pilot and staring at renewal, this is the reality check to bring to the meeting.

What held up

The core claim survived: the businesses that reported an outcome — a booked appointment, an opened claim, a saved subscription — could defend their spend, and the ones still reporting deflection could not. Deflection rate turned out to be exactly as useless in month twelve as we said it was in month one. It goes up when you make the escape hatch harder to find, which is the opposite of a healthy signal.

The revenue-capture case held up best of all. When roughly a quarter of inbound calls in home services go unanswered (per Invoca’s industry benchmarks, vendor-published) and most missed callers never call back, the math that moved the P&L was never labor savings — it was the bookings that used to leak away after hours. A year of data made that sharper, not softer: teams that wired the agent to a real calendar and measured book rate against their own baseline had the cleanest ROI story to tell.

What also held: the outcome has to be one thing finance already understands. The deployments that struggled to renew were often the ones that defined their “outcome” as something soft — a “good conversation,” a resolved sentiment score — instead of a countable event with a dollar value attached. A booked appointment, an opened claim, a saved subscription: each of those has a price the business can already quote. The framework was never about inventing a new metric. It was about attaching cost to the outcome the business was going to be judged on anyway.

The part that surprised us

The single biggest driver of a good cost-per-outcome number wasn’t model quality or automation rate. It was whether the team had an honest baseline from before the deployment. The ones who never measured their pre-AI answer rate or book rate couldn’t prove anything a year later, no matter how well the agent performed.

What didn’t hold up

The launch-era cost-reduction headlines did not survive the year. Vendor pitches promised 60–80% cost reduction; the deployments that ran a full twelve months landed much closer to independent estimates. McKinsey’s analysis of enterprise automation puts realistic net savings in the 20–35% range once you net out escalations, partial failures, and the engineering to keep the thing good — and that is roughly where the honest deployments actually sat.

Nobody hit the headline number. The teams that planned against it missed their year-one budget; the teams that planned against 20–35% beat it.

The other thing that didn’t hold: the assumption that the automation rate you saw in week two would be the rate you saw in month twelve. Resolution rates for genuinely self-service issues remain modest across the industry — reflected in the same consumer wariness Zendesk keeps documenting year over year. The agents that improved did so because a human sampled transcripts every week and fed the misses back in, not because the technology drifted upward on its own.

And one softer casualty: the idea that the number would settle quickly. It didn’t. Most teams needed three to six months before their cost-per-outcome figure stopped moving enough to defend it to finance — the time it took to accumulate a clean quarter of post-deployment outcomes, work the escalation edge cases out of the flow, and stop confusing launch-week novelty traffic with steady state. The framework rewarded patience, and punished anyone who quoted a month-one number as if it were durable.

The range operators should plan against

Here is the honest planning range after a year of watching deployments mature. These are directional bands, not promises — your number depends on your baseline answer rate, your outcome value, and how disciplined your transcript review is.

20–35%
realistic net cost reduction at 12 months (McKinsey range)
~25%
inbound home-services calls unanswered before deployment (Invoca)
3–6 mo
typical time to a defensible cost-per-outcome number
LeverLaunch-deck claim12-month reality
Net cost reduction60–80%20–35%
Automation / resolution rateflat, high, day oneclimbs only with weekly review
Payback driverlabor savingscaptured revenue first, labor second
Deflection ratethe headline KPIabandoned — gameable, not a business metric
Time to defensible ROI“immediate”3–6 months with a real baseline
Directional planning bands from mature (12-month) deployments. Net savings range per McKinsey (2024); the rest are operator-side ranges you should validate against your own baseline, not vendor guarantees.

Notice that the payback driver flipped columns. A year in, the teams with the strongest numbers weren’t the ones who cut the most headcount — they were the ones who stopped leaking revenue at 2 a.m. That tracks with the old, durable finding that small retention and capture gains compound into outsized profit, the kind of effect Bain has documented for years.

How to model your version of it

The framework hasn’t changed; the discipline around it has. To get a number you can defend at renewal:

  • Fix the baseline first. Answer rate, book rate, and revenue per booked outcome from before the agent existed. No baseline, no ROI story — that was the number-one failure mode this year.
  • Count the true cost. Platform, plus the human time on escalations and weekly transcript review. Leaving out the review labor is how teams talk themselves into the 60–80% fantasy.
  • Divide by outcomes, not contacts. Cost per booked appointment, opened claim, or saved subscription — the thing finance can price a service line against.
  • Plan against 20–35%, then beat it. Underwrite the case on the sober band and let captured revenue be the upside, not the load-bearing assumption.

What twelve months actually taught us

The lesson of the year isn’t that AI CX underdelivered. It’s that the teams who wrote down an honest number early were the ones still confident at renewal, and the teams who bought the headline spent month eleven explaining a gap. The framework works. It just rewards the operators who measure like they mean it.

Sources

Keep reading