Guides
How to measure the ROI of AI
Miguel Delgado · · 6 min read
How do you actually prove that the money you are pouring into AI is worth it, when you cannot see which teams are using it, for what, or whether a task costing you frontier-model prices could have run on something a tenth of the cost?
It is the question almost every leadership team is now circling, and very few can answer with numbers. Curative’s CEO put the problem bluntly on Harry Stebbings’ podcast: “The challenge now isn’t finding new use cases, it’s deciding when the spending becomes too large to justify.” Their Anthropic bill had grown sharply, month after month, as AI became embedded in more and more of the business. For a well-funded company still hunting for its edge, that is probably fine. For everyone else, that sentence is quietly terrifying, because “too large to justify” is a judgement, and almost nobody has the data to make it.
This is the strange place the market has arrived at. Adoption is no longer the hard part. The hard part is knowing whether any of it is paying off, and the honest answer for most organisations is that they simply cannot tell.
Why ROI is so hard to measure here
Traditional software has a tidy cost story. You buy a certain number of seats, you pay a predictable amount, and the bill next month looks like the bill this month. AI does not behave that way. Costs are usage-based, they scale with adoption, and they are driven by thousands of individual tasks running across teams that rarely talk to each other about how they are using the tool.
That creates two problems at once. The first is visibility: a single monthly invoice tells you the number went up, but nothing about where the money went or why. The second is quality: even if you could see the spend, you would still need to know whether the expensive work was actually worth doing expensively. An AI task can look productive and still be wildly overpriced, because the model doing it was far stronger, and far more costly, than the job required.
Put those together and you get the situation Curative described. Spend rises, value is assumed rather than measured, and eventually someone in finance asks a question nobody can answer.
A framework for measuring AI ROI
The good news is that this is a measurement problem, and measurement problems are solvable. Here is a practical way to work out whether your AI spend is justified, and where it is not.
1. Stop looking at the total. The aggregate bill is the least useful number you have. It tells you AI is being used and that it costs money, both of which you already knew. Real insight starts when you break spend down by team, by person and by workflow, because “is this worth it?” is only a meaningful question at the level of a specific task done by a specific team. Two teams with identical bills can be delivering wildly different value, and the total will never show you that.
2. See what the AI is actually being used for. Most organisations genuinely cannot say which tasks are consuming the majority of their AI budget. When you can see which workflows run most often, which teams are driving the load, and which surfaces the requests are coming through, the wasteful patterns become obvious almost immediately. The classic one is high-volume, low-stakes work — summarising a ticket, drafting a routine reply, tidying up notes — quietly running on your most powerful and most expensive model simply because nobody ever told it to do otherwise. Use a platform that can slice usage by skill, team and person, so this picture is something you can look at rather than something you have to guess.
3. Match the model to the task. Once you can see the load, right-sizing is straightforward. The routine 80% of work can move to whatever model is cheapest and good enough this month, while the frontier models stay reserved for the 20% of genuinely judgement-heavy work that actually needs them. Done well, the output quality does not drop at all. The only thing that changes is the bill. The trick is that you cannot make this call safely without the visibility from step two, because moving the wrong task to a cheaper model is exactly how quality falls off a cliff.
4. Cap the runaways. Some workflows misbehave. They loop more than they need to, reload their entire context on every call, or quietly consume far more than their output is ever worth. A quota or spend cap per team or per workflow catches these before they turn into a line item someone has to explain in a board meeting. Guardrails are not about using AI less; they are about making sure no single process can run away with the budget unnoticed.
From gut feel to a number you can defend
Notice what this framework does not ask you to do. It does not ask you to slow down adoption, ration access, or treat AI as a cost to be minimised. The companies getting this right are not the ones spending the least. They are the ones who can see exactly what they are spending it on, and can therefore make deliberate choices instead of blunt ones.
That is the real shift. “Too large to justify” is only frightening when it is a feeling. Once you can break your spend down by team and workflow, see which tasks are driving it, route work to the right model, and cap the outliers, that same phrase becomes a number, and a number is something you can stand behind when finance comes asking.
So it is worth asking yourself honestly: could you say right now which of your AI workflows is your most expensive, and whether it deserves to be? Most leaders cannot, yet. The ones who can are the ones who will still be scaling AI confidently long after everyone else has started nervously eyeing the invoice.
How Harriet can help you measure the ROI of AI
Everything above is easy to agree with and hard to do without the right tooling, which is exactly the gap Harriet is built to close.
Harriet sits between your teams and the models they use, which means every task runs through one place where it can be seen, priced and controlled. Its analytics slice usage by skill, team, person and surface, so you can see at a glance which workflows are driving your spend and which teams are behind them, rather than staring at a single opaque bill. That is step one and step two of the framework handled, without anyone having to assemble a spreadsheet by hand.
From there, the same platform lets you act on what you see. You can decide which teams get frontier models and which run on faster, cheaper ones, route routine high-volume work to the most cost-effective option automatically while reserving the best models for the work that genuinely needs them, and set quotas and hard cost controls per team or workflow so nothing runs away with the budget. Because Harriet is model-agnostic, you are never locked in: when a cheaper model becomes good enough for a given task, you move that task across and the saving lands immediately.
The result is that “too large to justify” stops being a feeling and becomes a number you can defend. You can show, task by task and team by team, what your AI is doing, what it costs, and why it is worth it. If that is the level of visibility you have been missing, we would love to show you how it works.
Common questions
How do you measure the ROI of AI?
Stop looking at the total bill and break spend down by team, person and workflow, because 'is this worth it?' is only meaningful at the level of a specific task. Then see what the AI is actually used for, match each model to the task it's doing (routine work on cheap models, judgement-heavy work on frontier ones), and cap the runaway workflows. Measured this way, 'too large to justify' stops being a feeling and becomes a number you can defend.
Why is AI ROI so hard to measure?
Unlike traditional per-seat software, AI is usage-based: costs scale with adoption and are driven by thousands of individual tasks across teams that rarely compare notes. A single monthly invoice tells you the number went up but nothing about where the money went, or whether the expensive work was worth doing expensively. That's both a visibility problem and a quality problem at once.
How do you cut AI costs without hurting quality?
Right-size the model to the task. The routine ~80% of work can run on whatever model is cheapest and good enough, while frontier models stay reserved for the ~20% of genuinely judgement-heavy work. Done well, output quality doesn't drop — only the bill does. The prerequisite is visibility: you can't safely move a task to a cheaper model until you can see what that task is and how it's performing.