How to measure the ROI of an AI implementation
- Reading time
- 11 minutes
- Assumes
- You've shipped or are about to ship an AI feature
- Updated
- Sep 6, 2026
Why the spreadsheet doesn't convince anyone
The standard artifact is hours saved per week, times a loaded hourly rate, times fifty-two. It produces a large number and it convinces nobody who wasn't already convinced.
Three reasons, and they're all fair.
Hours saved is not money saved unless somebody left or something else got done instead. If the same team does the same work in less time, you created capacity. That's real, and it isn't a line in the P&L.
The counterfactual is missing. The number compares against a world where nothing else changed, and that world doesn't exist. Volumes moved, a competitor launched, you hired two people, the process got cleaned up on the way.
The estimate came after the result. Everyone in the room knows the analysis was built once the outcome was visible, which means it could have been built to reach a different conclusion. This is the one that actually sinks the meeting, and no amount of analytical care fixes it afterwards.
So the fixes are procedural rather than analytical. The useful ones all happen before you ship.
What to write down before you ship
Here is the whole recommendation. One page, written before the change goes out, settled on a date you named in advance.
| Line | What goes in it |
|---|---|
| The change | One sentence on what you shipped, and when it landed. |
| What you expect to move | A metric, a direction, a number, a date. |
| What it shouldn't cost | A metric that has to stay within a stated tolerance of where it already was. |
| What would make you stop | The result that ends the work rather than explaining it. |
| What else is happening | The other things that could move these numbers, named now. |
| What it costs to run | Inference at production volume, review time, ongoing evaluation. |
That's it. It takes an afternoon, most of which is getting the before number out of a system rather than out of somebody's head.
Everything below is how to fill in each line without fooling yourself.
Why before and not after
Every line on that page is one you could write afterwards. Written afterwards, each one is a choice you made knowing the answer — which is exactly what the person across the table is discounting for. The page isn't more rigorous than a retrospective analysis. It's written before you knew the answer, and that's the part a skeptical reader can check.
Claim capacity unless headcount actually moved
The first decision is what unit you're claiming in, and it's worth taking a position on: claim capacity unless somebody actually left or a hire was actually avoided.
Three units are available.
Cost. Fewer hours, lower cost per case, lower vendor spend. Cleanest to measure and easiest to defend, and usually the smallest number.
Capacity. The same people doing more. Honest and easy to evidence, and it cashes out only if the extra throughput was something you wanted — a support team handling 30% more tickets matters if ticket volume was the constraint, and doesn't if it wasn't.
Quality or risk. Fewer errors, faster response, fewer escalations. Hardest to put in currency, and often the actual reason the work was funded.
The reason to default to capacity is that the conversion from hours to dollars is where these analyses lose the room. Multiplying saved hours by a loaded rate asserts that the hours turned into money, and if headcount didn't change, they didn't. A reader who works in finance spots that in about four seconds, and once they've spotted it they discount everything else on the page.
Overclaimed
The assistant saved us $420,000 a year in support costs.
Defensible
The same nine agents handled 31% more tickets. No roles were cut, so nothing left the P&L.
The second one is smaller and survives the meeting. If the capacity later absorbs growth you'd otherwise have hired for, that's when it becomes a dollar figure — and you'll be able to point at the hire that didn't happen.
Unsettleable
Improve support efficiency.
Settleable
Median handle time at or under 6 minutes by 31 March, from 8m12s across the 28 days before launch.
A claim has to be checkable by somebody who wasn't in the room. A metric, a direction, a number and a date is the smallest thing that qualifies.
If the work produced all three, report all three separately. One blended figure with the assumptions buried inside it reads as advocacy no matter how careful the arithmetic was.
Say what it shouldn't cost, too
The single most common gap in these pages is that they only say what should get better. Write a second kind of claim: what this shouldn't cost.
What you expect to move is the reason you did the work. A metric, a direction, a target, a date.
What it shouldn't cost is the thing you'd be embarrassed to break on the way there. It isn't a target — it's a tolerance around where the number already was. Pick the one a skeptic would go looking for. If you sped up responses, the guardrail is quality. If you automated a decision, it's the error rate on the cases the automation now handles alone.
Vague
Quality shouldn't suffer.
Checkable
CSAT shouldn't drop more than 2% below its average over the 28 days before launch.
A tolerance and a baseline window make a guardrail settleable. Without both, it gets waved through at the review by whoever is most invested in the result.
Keep the two kinds of claim separate when you settle them, because they aren't worth the same. Reaching a target is an achievement; holding a line is the absence of a loss. Score them together and "the change did nothing but broke nothing" comes out looking exactly like "the change worked but cost us something" — and those should lead to opposite decisions.
Get the before number, and its normal swing
This is the step that decides whether the whole exercise works, and it's the one teams skip.
Get the number from the system. Not an estimate from the team lead, who will be wrong by more than the effect you're trying to detect. Write down the query, the window, and the definition.
Write the definition down in painful detail. "Handle time" means about six things and your team uses three of them. The definition drifting between the before measurement and the after measurement is the most common way these analyses die quietly.
Get its normal swing too, and this is the part almost nobody does. Pull enough prior periods to see how much the metric moves on its own when nothing is happening. Then apply the test: if the metric's ordinary week-to-week variation is as large as the effect you're expecting, you will not be able to detect the effect. A number that swings 15% on a quiet week cannot demonstrate a 10% improvement, and no analysis performed afterwards will rescue it.
That test has three possible outcomes and all of them are useful before you build. Lengthen the window so the noise averages out. Pick a less jumpy metric. Or accept that this particular claim isn't provable and pick a different one while it's still cheap to change your mind.
Make the attribution argument explicit
You will not get a clean causal answer without a controlled experiment, and most of this work can't be cleanly experimented on. What you can do is make the attribution argument something a reader can inspect rather than something buried in a footnote.
Hold out a slice. One team, one region, one segment left on the old process for a month. It is the single highest-value thing on this page, it's routinely rejected as too complicated, and it's routinely regretted about six weeks later when somebody asks how you know. A partial holdout beats no holdout by an enormous margin, and it beats any amount of after-the-fact adjustment.
Compare against where the metric was heading. If it was already improving before you shipped, that improvement isn't yours. Take the shape of the pre-period and compare against where it was heading, rather than against where it sat.
Name the confounders in writing, before the result. Seasonality, a pricing change, a launch, a headcount change, the process cleanup you did at the same time. Written in advance they're context. Written afterwards they're excuses.
Report a range. A point estimate invites an argument about the point. A range with its assumptions stated invites a conversation about the assumptions, which is the conversation worth having. Bound it by the two honest extremes: the effect if none of your named confounders mattered, and the effect if all of them did.
Common mistake
Attributing the whole change to the AI when the project also cleaned up the process, rewrote the templates, and retrained the team. Those are often the larger effect. If you shipped them together you cannot separate them — so claim the bundle and say so. A credible claim about four things is worth more than an unprovable one about a single component.
Count the costs people leave out
The denominator is where these analyses lose credibility fastest, because the omissions always run in the same direction.
- Inference at production volume, not at pilot volume. This is the one that moves by an order of magnitude between the demo and the rollout.
- Build time at a real loaded rate, including the people who weren't on the project full time.
- Ongoing evaluation and maintenance. It isn't zero and it doesn't end. Models change underneath you and the thing has to be re-checked.
- Human review time. A step that pays a person to check every output has a running cost, and it belongs in the model rather than in a footnote.
- The cost of being wrong, at whatever rate you're actually wrong — which you should know, because you measured it.
A figure with a credible cost side persuades at a lower number than an inflated one with a thin denominator. The person you're presenting to is reading for the omission, and finding none is worth more than any amount of upside.
Settle it, including when it's bad
On the date you named, compare against what you wrote and record the outcome. Met, missed, or mixed.
This only compounds if the misses get recorded. A team with twenty settled AI decisions, six of which missed, can tell you which kinds of work pay off here. A team with twenty successes has a marketing document.
That's the real return, and it arrives later than the project does. The question a finance team is eventually asking isn't "did this project work" but "does this kind of project work here, and how much should I believe the next estimate." No single project answers that. A run of settled predictions does, and it gets more useful every time one of them turns out wrong.
Before you present a number
0 of 8 checked