Attribution for Conversations: Proving Which Messages Made Money
Conversational attribution is genuinely hard, and most vendors overclaim. Here is how to use holdouts and incrementality testing to prove which messages actually moved revenue — and how to report it without lying to yourself.
Every dashboard that promises to tell you exactly which text message made the sale is, on some level, lying to you. Not maliciously — the math is just harder than the interface admits. A customer sees your SMS, ignores it, gets a voicemail, searches your name, and books three days later. Five tools will each claim full credit for that booking. Conversational attribution done honestly means giving up the fantasy of perfect credit and learning to prove incremental revenue instead.
This is a technical guide for the person who has to defend a marketing number to a CFO. The argument is narrow on purpose: stop trying to trace every conversation to a dollar, and start measuring the lift your conversations create over a world where you sent nothing. That single shift — from tracing to testing — is what separates a defensible ROI story from a hopeful one.
Why conversational attribution breaks
Web attribution was never great, but it at least had a click to anchor on. Conversations don’t. An SMS gets read on a lock screen with no click. A voice agent has no UTM. A reply comes in over a different channel than the message went out on. The result is that the two default models both fail in opposite directions.
- Last-touchhands 100% of the credit to whatever happened right before the sale — usually the branded search or the “confirm your appointment” text — and makes your best top-of-funnel conversations look worthless.
- Multi-touch spreads credit across every touch with a tidy-looking formula, but the weights are invented. There is no ground truth that says the first SMS deserved 40% and the reminder 20%; someone picked those numbers.
Both models share the same original sin: they assume the sale would not have happened without the touch they are crediting. That assumption is exactly the thing you have not measured. Privacy changes have made it worse — Apple’s Mail Privacy Protection alone broke open-tracking as a reliable signal, and Litmus has reported that Apple Mail opens make up the majority of tracked opens (vendor-published), which means a huge share of “engagement” data is now noise dressed as signal.
The takeaway
Incrementality is the only honest answer
The clean way to prove a message made money is to compare people who got it against otherwise-identical people who didn’t. The gap between the two groups is the lift — the revenue that exists because of the program, not merely alongside it. This is the same logic clinical trials use, and it is the only method that survives a skeptical finance review.
The reason vendors avoid it is unflattering: incrementality almost always reveals that the credited number was inflated. When Meta’s own researchers compared standard attribution against randomized experiments across a large set of campaigns, they found attribution routinely overstated the true causal effect — sometimes by a wide margin (per Facebook’s published research on attribution vs. randomized experiments, vendor-published). The pattern generalizes: any model that credits correlation will flatter you, because a lot of the people you messaged were going to convert regardless.
If you can’t construct a group you didn’t message, you can’t prove the messages worked. You can only assert it.
How to run a holdout you can defend
A holdout is a randomly selected slice of your audience that is deliberately excluded from a campaign so it can serve as a control. The mechanics are simple; the discipline is not.
- Randomize at the person level, before the send. Split your eligible audience into treatment and control by a coin flip, not by who happened to have a phone number on file. Freeze the groups before anyone gets a message.
- Size the holdout to the outcome, not the audience.A 5% holdout on a list of 2,000 leaves you far too few conversions to detect a real difference. If your booking rate is low, you need a bigger control or a longer window — otherwise you will “measure” noise.
- Measure the business outcome, not the message metric.Count booked appointments, closed jobs, or revenue — the thing the CFO cares about — in both groups over the same window. Delivery and reply rates are diagnostics, not the result.
- Report the difference as a range with a confidence interval.The honest output is “the program drove an estimated 6–11% lift in bookings,” not “the program drove $48,213.” The range is the truth; the single number is theater.
- Keep an always-on holdout for the program as a whole.A permanent 5–10% global control tells you what your entire conversational channel is worth, campaign by campaign, without re-running the setup each time.
The uncomfortable part is that holdouts cost you something real: the revenue you would have earned from the people you deliberately did not message. That is the price of knowing. A 5% holdout on a healthy program is a rounding error against the cost of scaling a channel that turns out not to work.
What honest numbers actually look like
Set expectations before you run the test, because the credited number and the incremental number will not match, and someone will ask why. Well-run conversational programs do produce real lift — the point is that “real” is smaller and more defensible than the dashboard headline.
Speed is worth calling out because it is one of the few levers with strong independent evidence behind it. The classic Harvard Business Review lead-response research found that contacting a web lead within about five minutes dramatically raised the odds of qualifying it versus waiting even 30 minutes — a causal story you can actually test with a holdout, rather than assume.
| Question | Attribution dashboard | Incrementality (holdout) |
|---|---|---|
| What it measures | Touches near the sale | Sales that wouldn't exist otherwise |
| Typical result | One precise-looking dollar figure | A lift range with a confidence interval |
| Failure mode | Credits sales that would've happened anyway | Costs you the control group's revenue |
| CFO reaction | "How do you know?" | "Show me the two groups." |
Reporting it without lying to yourself
Once you have a defensible number, the temptation is to dress it back up. Resist it. The habits that keep an attribution story credible over time are boring on purpose.
- State the model out loud.“This is a holdout-based lift estimate, not attributed revenue” belongs on the slide, not in a footnote.
- Never blend the two numbers.Reporting attributed revenue one quarter and incremental lift the next — and calling both “ROI” — is how you lose the room permanently.
- Show the range and the assumptions.Name the window, the holdout size, and what you counted as a conversion. A number you can reproduce is worth more than a bigger number you can’t.
Sources
Keep reading
