Economics and buying Playbook

How to Calculate AI ROI Before You Commit Budget

An AI business case is a forecast built from a handful of measured inputs: what the work costs today, how much of it the system will really handle, what it costs to build and run, and what else improves. Measure each one on your own operation, fund on the downside case, and the numbers will survive a CFO's review.

For CFOs, sponsors and CTOs deciding whether an AI system has earned its budget before the build begins.

Published
Reviewed
Reading time
17 min

The short answer

Calculate AI ROI from measured inputs: today's cost per case from your systems of record, an automation rate proven on your own graded historical cases, and the full build and running cost, including inference, hosting, evaluation upkeep, human review of exceptions and model migration. Add hard benefits beyond labor, compute payback and NPV for base, downside and upside cases, and fund only if the downside pays back.

Key takeaways

  • Take every baseline figure from a system of record: volume, handling time, error rate, external spend and a fully loaded labor rate. A baseline built from interviews is the first thing finance discounts.
  • Set the automation rate by grading a few hundred of your own historical cases and running the system against them. Vendor containment figures measure someone else's work with someone else's definition.
  • Count running cost in full: inference, hosting, evaluation upkeep, human review of exceptions and forced model migrations. About one in five respondents to McKinsey's 2026 survey say AI operating costs have constrained their AI use.1
  • Fund on the downside column. Run base, downside and upside cases with the same variables, and compare the break-even automation rate with the rate your graded cases showed.
  • After launch, report cost per resolved case against a holdout handled the old way. AI arrives alongside other changes, and without a holdout its contribution is hard to isolate.2

AI business cases fail in two familiar ways. One multiplies an optimistic hourly rate by optimistic hours and produces a figure nobody in finance believes. The other declines to estimate, and the project goes unfunded or is funded on faith. The method below measures every input on your own operation, counts every running cost and judges the case on its downside, the column a CFO reads first.

  • 37%of respondents attribute any EBIT impact to AI, essentially unchanged from 20251
  • 2 to 4 yrsfor a typical AI use case to reach satisfactory ROI, against 7 to 12 months expected of other technology2
  • 1 in 5respondents say AI operating costs, including token costs, have constrained their AI use1

The AI ROI formula

AI ROI is the net benefit a system produces over a set horizon, minus the cost of building it, divided by everything it cost to build and run. Seven measured inputs feed it, plus the discount rate and horizon that finance already sets.

Inputs, per month unless stated
  V    cases handled                       from the system of record
  C0   fully loaded cost per case today    labor, rework, external spend
  a    automation rate                     share resolved correctly, no human touch
  e    effort per exception                vs today's average case (1.0 = same)
  R    run cost                            inference, hosting, evaluation, sampled review, migration
  G    other hard benefits                 errors avoided, external spend replaced, revenue
  B    one-off build and change cost
  r    monthly discount rate               your annual hurdle rate, converted to monthly
  T    horizon in months

Cost after launch    = R + V × (1 - a) × e × C0
Net benefit  N       = V × C0 - cost after launch + G
Payback (months)     = B ÷ N               with a ramp: first month cumulative N exceeds B
NPV                  = -B + Σ (t = 1..T) N_t ÷ (1 + r)^t
ROI over horizon     = (Σ N_t - B) ÷ (B + Σ R_t)
The whole model. Every number in it comes from your own operation or from your finance team.

Step 1: Measure the baseline from systems of record

The baseline is today's fully loaded cost per case multiplied by volume, taken from the systems that already record the work. Pull volume from the ticketing, case management or ERP system over a full business cycle, so month-end and seasonal peaks are included. Take handling time from system timestamps, using time studies only where none exist, and the error and rework rate from QA samples, reopened tickets, credit notes or journal corrections. Add external spend on the same work: outsourced processing, agency fees and contractor hours.

Load the labor rate properly. In US private industry, benefits made up 30.0% of employer compensation costs in June 2026, so an hour of a person's time costs roughly the wage divided by 0.70, before any overhead for space, systems and supervision.3 Use your own payroll figures where finance has them. A bare wage rate leaves out at least 30% of what the hour costs.

Loaded hourly cost  = hourly wage ÷ wage share of compensation + overhead per hour
C0                  = handling hours × loaded hourly cost
                    + rework rate × rework hours × loaded hourly cost
                    + external spend per case
Baseline per month  = V × C0

Name the system and date range behind every line, so finance can audit the baseline and every report after launch can be compared with it.

Step 2: Set the automation rate from your own graded cases

The automation rate is the share of cases the system resolves correctly, end to end, with no human touch, measured on your own historical cases. Pull a few hundred real cases, stratified by case type and including the awkward ones, and have your best people record the correct outcome for each. Run a prototype against that graded set before the full build is approved, then in shadow mode on live traffic, and sort every case into three buckets: resolved correctly, resolved after a human edit, and handed to a person. Weight each case type's rate by its share of volume to get a.

Resolution and containment are different measures. Resolution is the customer's problem solved, judged by the QA standard your people are held to. Containment is a conversation that ended without a handoff, which many product dashboards report; a contained case that returns tomorrow as a complaint counts as a success there and a failure on your ledger. Vendor figures describe someone else's case mix and belong nowhere in your model.

Published evidence argues for caution. In Stanford's 2026 AI Index, agents' success on OSWorld, a benchmark of real computer tasks, rose from 12% to about 66%, and agents still failed roughly one attempt in three on structured benchmarks.4 Executives told Deloitte that proofs of concept on dummy data breed optimism, and problems surface once real data arrives.2

Step 3: Count the full cost to build and run

The full cost is the one-off build plus every recurring line needed to keep the system accurate, and the recurring lines are where business cases go wrong.

Cost lineWhen it fallsWhat drives itUsually missed?
Build: integration, graded case set, guardrails, testingOne-offSystems touched, action classes, data preparationRarely
Change management and trainingOne-off, front-loadedPeople affected, workflow redesignOften
InferencePer caseCalls per case × tokens per call × price, plus retriesUnderestimated
Hosting, logging and tracingMonthlyYour environment, trace retention, search indexesOften
Evaluation upkeepMonthlyNew case types, re-runs on every changeAlmost always
Human review of exceptions and samplesPer caseException volume × effort, plus QA sampling of automated casesAlmost always
Model migrationEach time a model is retiredProvider retirement notices, regression runs, prompt reworkAlmost always
Security, audit and complianceMonthly and annualRegulations in scope, access reviews, record retentionOften
Recurring lines go into R, except human handling of exceptions, which enters through (1 − a) × e × C0.

Inference needs its own estimate, because its unit price and its bill move in opposite directions. Stanford's 2025 AI Index found the inference cost of a system performing at a fixed benchmark level fell more than 280-fold between November 2022 and October 2024.5 Yet about one in five respondents to McKinsey's 2026 survey said AI operating costs, including token costs, had constrained their AI use.1 Agents make several model calls per case, retry and carry long context, so tokens per case can outgrow the fall in price. Estimate inference per case from shadow-mode traces, never from a price sheet; our guide to governing AI inference spend covers the controls.

Model migration is a certain cost that most cases leave out. One major cloud's published lifecycle policy gives most hosted models six months' notice before retirement and some only 45 days; after the end-of-life date, requests to the model fail, and migration does not happen automatically.6 Each forced migration means re-running the graded set, reworking prompts and re-approving the release, so budget at least one inside any horizon longer than a year. The build figure has its own drivers, covered in what drives AI agent development cost.

Step 4: Value the benefits beyond labor

Benefits beyond labor often decide the case, and each needs its own measured formula. Only hard benefits, the ones that change a budget line, belong in the base case; soft and strategic benefits go in the upside column.

  • Labor. Hours returned count as savings only where a budget line moves: overtime, contractors, outsourced processing, or hires avoided as volume grows. Hours freed inside a salaried team are capacity until someone decides what it is for.
  • Errors. Errors avoided per month × cost per error, where cost per error is the rework, credit notes, write-offs and penalties your ledger already records.
  • Speed. Cycle time converted to cash: early-payment discounts, lower days sales outstanding, quotes answered while the buyer is still deciding.
  • Revenue. Conversion or retention lift measured against a holdout group the system does not touch. Assumed lift belongs in the upside column.
  • Risk. Expected loss avoided, as probability × impact, for fraud, compliance findings or missed deadlines, net of the new risks the system introduces.
  • External spend. Outsourced processing and agency fees the system replaces, where MIT NANDA documented some of the largest back-office savings, without material cuts to internal staff.7

The organizations that report material returns treat this step as a redesign of the work. McKinsey's AI high performers, about 6% of respondents, attribute at least 5% of EBIT to AI; they pursue growth alongside efficiency, and nearly three-quarters of them have fundamentally redesigned workflows because of AI, against one-quarter of everyone else.1 A system inserted into an unchanged process returns hours; a redesigned process returns money.

Step 5: Compute payback and NPV

Payback is the build and change cost divided by the monthly net benefit, counted from the month benefits start; NPV discounts each month's net benefit at your hurdle rate and subtracts the build. Payback tells a CFO how long capital is at risk, and NPV whether the project beats the next use of the money.

Model a ramp, because benefits start after the build and grow as autonomy widens. Use the discount rate finance applies to other capital projects. Cap the horizon at three years, since the model, the prompts and the process will have changed by then, and treat anything later as upside.

Deloitte's 2025 survey of 1,854 executives found most respondents reach satisfactory ROI on a typical AI use case in two to four years, against the seven to 12 months expected of other technology investments, and only 6% reported payback inside a year.2 Narrow processes pay back faster: in a Gartner survey of 160 senior finance leaders in early 2026, data extraction, payables and receivables automation, and report creation generally delivered returns within nine to 10 months.8 A focused process with a measured baseline should sit with the second group. A case that needs three years to pay back is a transformation program and should be argued as one.

Step 6: Run base, downside and upside cases

Run the same model three times with the same variables, changing only their values, and fund the project on the downside column. The table uses hypothetical round numbers on an index where today's process costs 100 a month; B stands for your own build and change cost.

Variable (example)DownsideBaseUpside
Today's process cost per month, V × C0100100100
Automation rate a, from graded cases50%70%80%
Effort per exception e, vs today's average case110%100%90%
Run cost per month R1286
Other hard benefits per month G046
One-off build and change cost1.5 × BBB
Process cost per month after launch673824
Net benefit per month N336682
Cost per case after launch, before the build (today = 1.00)0.670.380.24
Payback, as a multiple of the base case3.0×1.0×0.8×
Example figures only. Every column uses the same variables and the same formulas; only the values move.

Read the downside column first. With automation 20 points below the graded result, exceptions 10% harder, running cost up by half, no benefit beyond labor and the build 50% over, this case takes three times as long to pay back. A base case that repays its build in eight months has a 24-month downside, which a CFO can fund. A base case that needs 18 months has a downside of four and a half years; narrow the scope to the case types with the highest graded rates before funding it.

One more figure tells the sponsor how much room there is: the break-even automation rate, the value of a at which the project just pays back inside your target period P.

Break-even automation rate
  a* = 1 - (V × C0 - R + G - B ÷ P) ÷ (V × e × C0)

Compare a* with the rate your graded cases showed; the gap is the margin of safety. A case whose graded rate sits 30 points above break-even can absorb a disappointing launch, and one sitting five points above cannot. Where no graded evidence exists yet, apply a harsher test: halve the automation rate, double the build cost and add 50% to the run cost. If the project still pays back inside two years, optimistic assumptions are not carrying the argument.

How to calculate cost per outcome

Cost per outcome is the all-in monthly cost of the process after launch divided by the cases it resolved, and it is the one figure to set beside today's cost per case in every report.

Cost per resolved case, all-in   = (R + V × (1 - a) × e × C0 + B ÷ T) ÷ V
AI run cost per case it resolved = R ÷ (a × V)

The second line catches what cost per token hides. Inference spent on cases that end with a person still has to be paid for, so it lands on the cases the system did resolve. When the automation rate slips, AI cost per resolved case rises even if the token price falls. Put both lines on the monthly dashboard: the first tells finance whether the process got cheaper, the second tells engineering where the run cost goes.

What the research says about AI ROI, and why the studies disagree

The major studies agree that most organizations cannot yet show a financial return from AI, and they disagree on how many because each one measures a different thing.

StudySample and datesWhat counted as a returnFinding
MIT NANDA, The GenAI Divide (July 2025)153 leaders surveyed, 52 organizations interviewed, 300+ public initiatives; January to June 2025Generative AI deployed beyond pilot with measurable KPIs, six months after the pilot95% of organizations getting zero return7
McKinsey, The state of AI in 2026 (August 2026)1,719 participants in 97 nations; May 4 to June 8, 2026Any EBIT impact the respondent attributes to AI37% report some; about 6% attribute 5% or more of EBIT1
IBM CEO study (May 2025)2,000 CEOs in 33 countries and 24 industries; February to April 2025AI initiatives that delivered the ROI expected of them25% delivered it; 16% scaled enterprise-wide9
Deloitte, AI ROI (October 2025)1,854 executives in Europe and the Middle East; August to September 2025Time for a typical use case to reach satisfactory ROITwo to four years; 6% within a year2
Gartner, finance AI (September 2026)160 senior finance leaders; January to April 2026Time to expected value, by finance use caseNine to 10 months for data extraction, payables and receivables, and report creation8
Five studies of AI returns, with the definition each one used.
Exhibit 1The share reporting a return depends on the question asked
  • MIT NANDA 2025: organizations with a measurable return5%
  • McKinsey 2026: 5% or more of EBIT from AI6%
  • Deloitte 2025: payback within a year6%
  • IBM 2025: initiatives delivering expected ROI25%
  • McKinsey 2026: any EBIT impact from AI37%
The unit changes from organizations to initiatives to use cases, and the bar for a return changes with it. Sources: [7], [1], [2], [9]

Three differences explain the spread. The unit: MIT counted organizations, IBM initiatives, Deloitte a typical use case. The bar: MIT required measurable P&L impact six months after a pilot, while McKinsey counts any EBIT impact a respondent attributes to AI. The scope: MIT studied enterprise generative AI tools, and the other surveys cover AI broadly. MIT's authors flag their own limits: a self-selected sample, success definitions that varied, and a six-month window that may understate success for complex systems.7 McKinsey's figures are respondents' own attributions.1

Read together, the studies support two conclusions. The base rate is poor, and 64% of the CEOs IBM surveyed said the risk of falling behind drives some technology investment before its value is clear, so the burden of proof sits on your own measured inputs.9 And returns concentrate where a specific process was measured before and after: back-office work in MIT's data, finance automation in Gartner's, and redesigned workflows among McKinsey's high performers.781

Build, buy or wait: run the same model on each path

Decide between building, buying and waiting by running the identical model three times, changing only the inputs each path actually moves, and choosing the path with the best downside NPV.

Exhibit 2The same model, three sets of inputs

Build

A system built on your cases

  • B is highest: integration, graded set and guardrails
  • a is tuned to your case mix and can rise with each release
  • R includes the evaluation upkeep and migrations you run
  • Code, prompts and graded cases stay with you

Buy

A product with AI features

  • B is lower, though integration and change management remain
  • a is still measured on your graded cases; the product is tuned to an average customer
  • R includes subscription or per-resolution fees that rise with volume
  • The vendor controls model changes; add an exit cost

Wait

Revisit on a set date

  • B and R are zero until you start
  • N is zero too: the cost of waiting is the monthly net benefit forgone
  • Inference at a fixed capability keeps getting cheaper, so a later build may run for less
  • Right when every downside NPV is negative, or the data is not ready
Pick the path with the best downside NPV. Waiting is a legitimate answer, and it has a price: N for every month you wait. Source: [5]

The evidence on building against buying points both ways. In MIT NANDA's interview sample, tools bought from or co-developed with external partners reached deployment about 67% of the time, against about 33% for internal builds, and the authors caution that the gap may reflect the organizations more than the approach.7 McKinsey's 2026 survey found 32% of respondents had decided against buying at least one software product or feature because they could build it internally with agentic coding tools.1 Neither figure settles your decision; the model, run on your numbers for each path, does.

How to report AI ROI after launch

Report the same variables the business case used, measured the same way, every month, against a holdout of cases the system does not touch. Executives interviewed by Deloitte said AI usually arrives alongside data clean-ups, reorganizations and process changes, and one could produce only a ballpark estimate because AI's gains were hard to separate from those programs.2 A random slice of cases handled the old way moves with everything else, so the difference belongs to the system.

The monthly AI ROI report

  • Cost per resolved case, all-in, against today's cost per case and against the holdout.
  • Automation rate on live traffic, by case type, against the graded-set rate that justified the build.
  • Effort per exception, from timestamps in the exception queue.
  • Error rate from QA sampling of automated cases, scored to your people's standard.
  • Run cost by line: inference, hosting, evaluation, review, and any migration in the month.
  • Hard benefits booked, each tied to a ledger entry or a budget line that moved.
  • Position against the downside case, with the agreed trigger for a review.

Agree the review trigger before launch: if the live automation rate sits below the downside assumption for two consecutive months, the sponsor and finance narrow the scope, fix the exception path or stop. A project that can be stopped on evidence is far easier to fund.

In the systems we build, the graded case set and the baseline come before the architecture, because together they decide whether the build should happen at all, and once it does they become the evaluation suite and the ROI dashboard. That is where our AI development work starts.

Questions leaders ask

What is a good ROI for an AI project?

Judge an AI project on payback and on its downside case, since a single ROI percentage hides both. A focused process with a measured baseline should repay its build within the payback period finance sets for other systems; Gartner found common finance automations returning value in nine to 10 months.8 Deloitte found most organizations take two to four years on a typical use case, so a longer case needs a strategic argument as well as a financial one.2

How do you measure the ROI of generative AI?

Measure it per case: cost per resolved case after launch, compared with the same figure for a holdout handled the old way. Time saved counts only when a budget line changes. McKinsey's 2026 survey shows why: 80% of respondents said AI improved their own productivity, while 37% reported any EBIT impact.1 Individual gains stay out of the P&L until the workflow and the budget change around them.

How is AI agent ROI different from automation ROI?

The formula is the same, and three inputs behave differently. Agents make several model calls per case, so inference varies widely; their mistakes can be actions, so exceptions and approvals cost more; and autonomy widens in stages, so benefits ramp slowly. Deloitte found only 10% of agentic AI users seeing significant ROI so far.2 Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls.10

Is it true that 95% of AI projects fail?

The figure comes from MIT NANDA's 2025 report, which found 95% of organizations getting zero measurable return from generative AI.7 It rests on 153 survey responses, 52 interviews and a review of over 300 public initiatives, with impact measured six months after pilots, and its authors note that success definitions varied. McKinsey's 2026 survey, with a looser definition, found 37% of respondents attributing some EBIT impact to AI.1

Can I use an AI ROI calculator?

Yes, if you control its inputs. A calculator is the formulas in this article with defaults filled in, and the defaults decide the answer. Replace every default with a measured figure: volume and cost per case from your systems of record, the automation rate from your graded cases, and a run cost that includes inference, evaluation, review and migration. A vendor's calculator starts from the vendor's claims.

Which costs do AI business cases usually miss?

The recurring ones: evaluation upkeep as case types change, human review of exceptions and quality samples, hosting and trace retention, security and audit work, and forced migrations when a provider retires a model. One major cloud gives most hosted models six months' notice and some only 45 days.6 Change management and training, both one-off, are also routinely left out.

Sources

  1. The state of AI in 2026: On the road to ROIMcKinsey & Company, August 25, 2026
  2. AI ROI: The paradox of rising investment and elusive returnsDeloitte, October 22, 2025
  3. Employer Costs for Employee Compensation, June 2026U.S. Bureau of Labor Statistics, September 9, 2026
  4. The 2026 AI Index ReportStanford Institute for Human-Centered Artificial Intelligence, 2026
  5. The 2025 AI Index ReportStanford Institute for Human-Centered Artificial Intelligence, 2025
  6. Model lifecycleAmazon Web Services, Amazon Bedrock User Guide
  7. The GenAI Divide: State of AI in Business 2025MIT NANDA, MIT Media Lab, July 2025
  8. Gartner Says CFOs Must Take a More Disciplined Approach to Finance AI InvestmentGartner, September 24, 2026
  9. IBM Study: CEOs Double Down on AI While Navigating Enterprise HurdlesIBM, May 6, 2025
  10. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027Gartner, June 25, 2025

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.

Read next

All insights
  • On a plinth, a street of buildings whose heights are the shares of an AI agent's build effort: integration is the tallest tower, spanning out to each of your systems of record, while evaluation, agent logic, guardrails and the console step down beside it. The model is a small glass cube on its own pedestal at the end.

    AI agents Analysis

    What Drives AI Agent Development Cost in 2026

    For CFOs and CTOs deciding how much to commit to a first production AI agent, and what to settle before an RFP goes out.

    16 min read

  • Three AI agents send their calls through one lit router, which passes most of them to a fleet of small models and only a hard one to a large frontier model. Each agent has its own budget gauge; one has spent to its cap and a red barrier stops it, while the other two keep working.

    Economics and buying Playbook

    AI Inference Cost: How to Govern LLM and Agent Spend

    For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.

    16 min read

  • Your AI system stands as the tallest tower on your own ground, with its code, evaluation set, model weights and vector index beside it, each ticked as it lands. A bridge brings in the builder's work from its island, and a second bridge hands the system to another team that can run it.

    Economics and buying Checklist

    15 Questions to Ask an AI Development Company Before You Sign

    For CTOs, procurement leads and general counsel choosing between shortlisted AI development companies before a contract is signed.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.