Why the revenue question gives a bad answer
PwC put the question to 4,454 chief executives across 95 countries for its 29th Global CEO Survey. Fifty-six per cent said AI had produced neither higher revenue nor lower costs in the twelve months to autumn 2025. Twelve per cent reported both at once. The survey records the answers without explaining them, and our own reading is that much of that gap is measurement: the work started on a process nobody had timed.
Before any of this there is an earlier question, whether the process is worth automating at all, and when AI pays off and when it is too early is a separate decision. This piece assumes you have made it and now need a number.
The first instinct is to look at revenue. Sales went up after launch, so the system worked.
It does not survive contact with a finance team. In the same quarter you probably changed pricing, hired two people, ran a campaign and lost a competitor. Nobody can separate the model’s contribution from the rest, and any number you present will be a guess wearing a suit.
Cost reduction has the same problem one level down. Headcount often stays where it was, because people move to other work rather than leaving. The saving can be real and still never appear as a line in the P&L.
So the useful measurement is not financial at first. It is operational, and it lives inside one process.
The three numbers you need before anyone writes code
Measuring roi on ai projects starts with arithmetic that any operations lead can do on a napkin, using figures that exist before the project does.
- How long the task takes today. Not the estimate, the measured time. Someone sits with a stopwatch, or you pull timestamps out of the system that already records them.
- How many times a day it happens. Volume is what turns seconds into budget. A task that takes four minutes and runs four hundred times a day is a different business case from the same task running twice.
- What that person would do instead. This is the number people skip. If the freed hour goes to work that was already waiting, the saving is real. If it goes to nothing in particular, you have bought idle time.
Multiply the first two, apply the third, and you have a defensible figure before a single line of code exists. AI productivity gains only count once the freed hours land on work that was already waiting. Everything after that is comparing this number with what the build and the running costs come to.
The “before” number is the whole problem
Here is the part that decides whether the project can be measured at all.
If nobody can say how long the task takes today, you have no baseline, and without a baseline there is nothing to compare the result against. You will finish the project, the team will feel faster, and the finance side will ask for evidence you cannot produce. That is one of the quieter reasons pilots stall on the way to production: not that they failed, but that nobody could show they worked.
This is why the first process to automate is almost never the most exciting one. It is the one where the “before” figure is already known, or can be measured in a week. Choosing on that basis feels unambitious and pays for itself twice: you get a project that works and a number you can show.
If the baseline does not exist, measuring it is the first piece of work. It costs almost nothing and it is the difference between an outcome and an opinion.
The half of the cost that gets left out
An AI ROI calculation fails just as often on the cost side, because the cost of an AI system has two halves and proposals usually show one.
| What it covers | When you pay | |
|---|---|---|
| Build | Setting up the model, connecting systems, testing on real cases | Once |
| Run | Model calls, hosting, monitoring, keeping it current as the process changes | Every month, forever |
The running half is the one that surprises people. Every request to a model costs money at the moment it is used, so the bill grows with adoption. A system nobody uses is cheap to run and worthless. A system everyone uses is valuable and has a real monthly line next to it.
There is a second running cost that has nothing to do with the model. The process underneath it keeps changing: new products, new rules, new wording. Somebody has to keep the system current, and if nobody is assigned to that, accuracy drifts quietly until people stop trusting the output.
Put both halves against the saving from the section above and you have the actual return, rather than a comparison between a one-off invoice and an ongoing benefit.
A worked example
Numbers make this concrete, so here is the shape of the calculation. The figures below are illustrative rather than a client case.
Take a support team that reads every incoming ticket and routes it by hand. Someone times it: three minutes per ticket, four hundred tickets a day, handled by a team of three.
| Figure | |
|---|---|
| Time per ticket today | 3 minutes |
| Tickets per day | 400 |
| Hours spent per day | 20 |
| Share the system routes without a human | 70 per cent |
| Hours returned per day | 14 |
Fourteen hours a day looks strong until you apply the third question from earlier: what do those people do instead. If a backlog of unanswered second-line tickets has been growing for a year, the hours land somewhere useful and the case holds. If nothing is waiting, you have bought capacity you do not need, and the honest number is much smaller.
Then subtract the running cost. Every ticket the system touches is at least one model call, and the awkward ones are several, so the monthly bill scales with the same volume that produced the saving. That figure belongs on the same page as the hours.
None of this requires a model to exist yet. That is the point of working out AI ROI this way: the case is built from numbers you already have, and the project either clears the bar or it does not.
Which AI metrics survive a finance review
Not every number that looks good in a demo will hold up when someone sceptical reads the slide. These are the ones that tend to survive.
- Minutes per case, before and after. Measured the same way both times, by the same method, on the same kind of work.
- Volume handled without adding people. Clean, visible and hard to argue with, especially when demand is growing.
- Share of cases that finish without a human. The honest version also reports what happened to the remaining cases and who picked them up.
- Rework rate. If the system is fast and wrong, speed is a liability. This is the metric that keeps the others honest.
- Time to a correct answer. How long someone waits before the answer they receive is actually right. It is the only measure here a customer would notice without being told about it.
What does not survive: model accuracy on a test set, tokens processed, hours of engineering saved by a coding assistant. These are AI metrics that describe the system rather than the business, and a finance reviewer will say so.
Proving AI ROI to executives
The presentation problem is different from the measurement problem, and it catches teams that did the measurement properly.
Lead with the baseline and let the result follow it. A room that has not agreed on how long the task used to take will spend the whole meeting arguing about the denominator instead of making the decision it came for.
Give the range rather than the best case. A saving quoted as a band, with the volume assumption stated next to it, reads as competence. A single confident figure invites someone to find the assumption that breaks it, and there is always one.
Put the running cost on the same slide as the saving. Hiding it does not make it go away, it makes the second quarter an unpleasant conversation.
And say what you did not measure. Every honest estimate has an edge, and naming it yourself costs you less than having it found for you.
When the honest answer is to wait
The AI value on offer is not always positive, and a vendor who never says so is not being useful.
Wait if the process changes every few weeks. You will spend the budget chasing the process rather than improving it.
Wait if the volume is low. Four minutes saved on a task that runs twice a day comes to roughly thirty-five hours across a working year. Real, and still not worth a project.
Wait if the knowledge the system needs lives in people’s heads and nowhere else. Writing it down is valuable on its own, and it has to happen first regardless.
We say this on review calls more often than clients expect. The alternative is a project that works technically and cannot justify its own line in the budget, which is a worse outcome for everyone.
What next
If you are trying to put a number on a project that has not started yet, bring the process to a review call. It is free and takes fifteen to thirty minutes: we go through the task, the process around it and the data behind it, which is where the baseline and the running cost both come from. Sometimes the answer is that the volume does not justify the work, and that is a result worth having before the budget is committed.
- Why the revenue question gives a bad answer
- The three numbers you need before anyone writes code
- The “before” number is the whole problem
- The half of the cost that gets left out
- A worked example
- Which AI metrics survive a finance review
- Proving AI ROI to executives
- When the honest answer is to wait
- What next