Measuring ROI on enterprise AI
An AI program that cannot state its return in the language of the business will lose its budget. How to instrument outcomes so the value survives finance review.
An enterprise AI program that cannot state its return in the language of the business will eventually lose its budget — no matter how good the technology is.
The failure here is rarely the value. The systems often work. The failure is measurement: no one defined the outcome, instrumented the baseline, or attributed the improvement in a way finance could accept. The result is a program that feels valuable and cannot prove it.
Define the outcome before you build
The outcome metric is not "AI adoption" or "queries answered." It is the business number the system is supposed to move: average handling time, qualified-lead conversion, escape-defect rate, time to resolution. Name it first, in the units the business already tracks, and design the system to move it.
A defensible number beats an impressive one.
Instrument the baseline first
You cannot claim an improvement you did not measure. Before the system goes live, capture the baseline: what the metric is today, how it varies, what drives it. The most common reason an ROI claim collapses under scrutiny is that there was nothing to compare against.
Attribute conservatively
When the metric improves, resist the urge to claim all of it. Some of the change is seasonality, some is other initiatives, some is noise. A conservative attribution that survives a skeptical review is worth far more than an aggressive one that invites it. Where you can, run a holdout or a staged rollout so the comparison is clean.
Report cost per outcome
Finance does not care about cost per token. It cares about cost per resolved case, per qualified lead, per inspected unit. Tie the running cost of the system to the outcome it produces, and the conversation shifts from "what does the AI cost" to "what does each result cost, and is that a good trade." That is the language in which AI moves from a cost center to a line on the P&L.
Key takeaways
- Define the outcome metric before you build — handling time, conversion, defect rate, not "AI adoption".
- Instrument the baseline first; you cannot claim improvement you did not measure.
- Attribute value conservatively; a defensible number beats an impressive one.
- Tie cost to outcome — report cost per resolved case, not cost per token.
Start a conversation