How to measure value from enterprise AI
Why most AI value debates are really arguments about a missing baseline.
Measuring value from enterprise AI requires four things: a baseline established before deployment, a named owner for each claimed benefit, a measurement cadence that outlasts the implementation project, and provenance showing which AI outputs were actually acted on.
Three of those are ordinary value realization discipline. The fourth is specific to AI, and it is the one most organizations omit.
Why AI value arguments go in circles
An organization deploys an AI capability. Some months later leadership asks what it produced. Advocates cite time saved; sceptics point out that headcount did not change and the metric would have improved anyway. Neither side can settle it, because no counterfactual was captured — nobody recorded the starting position, and the question was not framed as measurable before the money was spent.
Most disputes about AI value are not disagreements about AI. They are the predictable result of deploying without a baseline, which makes the question unanswerable rather than merely contested.
What to establish before deployment
- The baseline. Current performance on the specific measures the investment is supposed to move, captured before anything changes.
- The benefit owner. A named person accountable for each claimed benefit, who remains accountable after the project closes.
- The measurement approach. Agreed in advance, including what would count as failure. An investment that cannot fail cannot be evaluated.
- The review cadence. Continuing past go-live, because the implementation programme disbanding is exactly when measurement stops.
The AI-specific requirement: provenance
For conventional technology, usage is a reasonable proxy for effect. For AI it is not: a system can generate large volumes of output that nobody acts on, and usage metrics will look healthy throughout.
What is needed instead is a record of which outputs entered a decision — which recommendations were approved, by whom, and what followed. That is a provenance requirement, and it is why EXOS attaches a chain of custody to every output. Without it, AI value measurement reduces to counting activity.
What to be sceptical of
- Time-saved figures with no baseline. Usually derived from a survey asking people to estimate hours.
- Aggregate productivity claims. Real effects are concentrated in specific tasks; aggregates hide whether anything changed where it mattered.
- Adoption as a proxy for value. Adoption measures whether people opened it.
- Benchmark comparisons. Another organization's result tells you little about yours, because the binding constraint is usually knowledge quality and process, not the model.
Where EXOS sits on this
Baselines, benefit ownership and tracking structure are the object of value realization. The provenance requirement is met through Tacit OS, which records what informed each output and who approved it.