Economics and buying Playbook
AI Inference Cost: How to Govern LLM and Agent Spend
Per-token prices keep falling and AI bills keep rising, because every task now consumes more tokens. The spend becomes governable when you manage cost per completed task, pull five levers behind an evaluation gate and put a budget in code around every agent.
For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.
The short answer
AI inference cost is the recurring spend on every model call a system makes in production. Govern it with one unit, cost per completed task: all model, retrieval, platform and review costs divided by the tasks that passed your quality check. Then route easy work to smaller models, cache repeated context, cap context and output, batch what can wait, and give every agent a budget with a hard stop.
Key takeaways
- Manage cost per completed task: everything the system spent in a month, including the time people spent reviewing its work, divided by the tasks that passed your quality check.
- Token prices fall and bills still rise because tokens per task grow. On Epoch AI's benchmarks, reasoning models' responses were lengthening about 5x a year.1
- Listed prices mislead. In a 2026 study, the model with the lower listed price cost more in total in 32% of model-pair comparisons.2 Test every candidate on your own tasks.
- Give every agent a budget enforced in code, per task, per day and per month, with a kill switch. OWASP lists unbounded consumption among the top ten risks for LLM applications.3
- Start with showback by team, product and agent, which the FinOps Foundation treats as required, and add chargeback only where your accounting policy calls for it.4
Building an AI system is a project with an end date. Running it is a bill that arrives every month and moves whenever a model, a prompt or user behavior changes. The build is covered in what drives AI agent development cost. This playbook covers the run: the unit to measure, where the tokens go, the levers that cut the bill and the budgets that cap it.
- 98%of State of FinOps respondents now manage AI spend, up from 31% in 20245
- 280x+fall in inference cost at a fixed capability level between November 2022 and October 20246
- 32%of model pairs in a 2026 study where the model with the lower listed price cost more in total2
Why AI bills rise while token prices fall
AI bills rise because the tokens each task consumes are growing faster than the price per token is falling. The price side is real: Stanford's 2025 AI Index found that the inference cost of a system matching a leading late-2022 chat model fell more than 280-fold between November 2022 and October 2024.6 Finance budgets for that trend. Engineering then adopts the newest capability, which spends far more tokens on the same work.
Three forces push tokens per task up. Reasoning models bill their hidden reasoning as output; on Epoch AI's benchmarks their responses were lengthening about 5x a year, against 2.2x for other models, and ran about 8x longer on average.1 Agents make many calls per task, and each call re-sends the instructions, tool definitions and history that came before it. Retries, loops and abandoned attempts consume tokens and produce nothing.
Listed prices are a poor guide to task cost. A 2026 study of eight frontier reasoning models across 12 tasks found that in 32% of model-pair comparisons the model with the lower listed price cost more in total, by up to 28 times, and that repeated runs of one query varied by up to 9.7x in reasoning tokens.2 Compare models on your own tasks at their real consumption, and forecast spend as a range.
Falling
Price per token at a fixed capability
- Inference at a fixed capability level: over 280-fold cheaper in two years
- Increasingly capable small models
- Hardware costs falling about 30% a year
- Open-weight models closing the gap with closed ones
Rising
Tokens consumed per completed task
- Reasoning responses lengthening about 5x a year on benchmarks
- Agent steps that re-send instructions, tools and history on every call
- Wider retrieval and raw tool outputs in the context
- Retries, loops and abandoned attempts
The run also outlasts the build. A FinOps Foundation working group paper notes that inference costs accumulate for as long as a model is in use, often reaching 80 to 90% of total cost of ownership.7 A budget process that scrutinizes the build and waves through the run is watching the smaller number.
Cost per completed task: the unit to manage
Cost per completed task is everything the system cost in a period divided by the tasks it finished to an acceptable standard in that period. It connects the AI bill to the business, because a CFO can set it beside what a task is worth: the cost of a claim handled by a person, or the margin on an order.
cost per completed task =
( model tokens: input + output + reasoning
+ retrieval: embeddings, vector search, reranking
+ tool and API calls the agent makes
+ platform: gateway, hosting, logging, tracing
+ human time: reviews, approvals, handled escalations
+ evaluation runs, spread across the month )
÷ tasks completed that passed the quality checkTwo choices keep that formula honest. The numerator includes people: an agent that escalates a third of its cases has moved cost into payroll, so a model swap that raises escalations can cut the token bill and raise the true cost. The denominator counts only tasks that passed, so a cheaper model that fails more often can cost more per completed task than the one it replaced.
The FinOps Foundation's AI guidance lists resource metrics such as cost per inference, cost per token and cost per API call.8 Engineers need those to find waste. Report upward in business units, and split every change in the bill into its causes.
| Cause | What changed | Owner | First response |
|---|---|---|---|
| Volume | More tasks attempted | Business owner | Update the forecast if value per task holds |
| Mix | More hard tasks reaching expensive models | Product owner | Review the routing rules |
| Tokens per task | Longer context, more steps, reasoning or retries | Engineering lead | Find the category that grew |
| Effective price | Provider price, contract, cache hit rate or hosting | FinOps and procurement | Renegotiate, recommit or fix the cache |
| Pass rate | Fewer tasks passing the quality check | Engineering and QA | Treat it as a quality incident first |
Where the tokens go in an LLM or agent task
In the systems we build, most of a task's tokens go to material the user never sees: instructions, retrieved documents, history, tool results and hidden reasoning. Each of the seven places below grows for its own reason and has its own first fix.
| Where tokens go | Why it grows | First lever |
|---|---|---|
| System prompt and tool definitions | Re-sent on every call; grows with each policy or tool added | Prompt caching; give each step only the tools it needs |
| Retrieved context | Twenty passages sent when five would answer | A context budget, with reranking |
| Conversation and agent history | Every step re-sends every earlier step | Summarize old turns; cache the stable prefix |
| Tool outputs | Raw API payloads pasted in whole | Return only the fields the next step needs |
| Reasoning tokens | Billed as output; reasoning models produce about 8x more tokens1 | Route simple tasks elsewhere; set a reasoning effort level |
| Final output | Verbose defaults and free-form prose | Output limits and structured formats |
| Retries and loops | Timeouts, invalid output, repeated failing steps | Step limits, validation before retry, a task budget |
Instrument before optimizing. Log token counts by category on every call, tagged with task, agent and step, in the OpenTelemetry traces your platform already collects. Teams that skip this trim the prompt they can see and miss the retry loop that doubles the bill.
How to reduce LLM costs without losing quality
Five levers do most of the work: route each task to the cheapest model that passes, cache what repeats, budget the context, batch what can wait and cap the output. Each changes what the model sees or does, so each ships behind your evaluation gate like any other release.
Route by difficulty
A single model sized for the hardest request overpays on every easy one, so a router sends each request to the smallest model likely to pass. In published research, routers trained on human preference data cut cost by more than half in some settings without reducing response quality on standard benchmarks.9 A FinOps Foundation paper adds that sending work that needs no reasoning to a non-reasoning model can save 4 to 20 times the tokens.7 In a September 2026 survey, 86% of large companies were evaluating or using a router.10 Test the routing rules against graded cases for each tier: a hard case sent to a small model trades a token saving for a failure.
Cache what repeats
Prompt caching lets the provider, or your own inference server, reuse the processed form of a prompt prefix it has seen recently, and bills cached input well below fresh input; the FOCUS working group notes the gap can be an order of magnitude.11 Across three major providers and more than 500 agent sessions with 10,000-token system prompts, a 2026 evaluation found caching cut API cost by 41 to 80% and time to first token by 13 to 31%.12 Layout decided the result: stable content first, dynamic content last, changing tool results outside the cached block. Caching everything by default was less consistent and sometimes slower.12
Semantic caching stores answers and returns one when a new question is close enough to an old one. Use it only where the answer is identical for every user, such as a policy question, with a strict similarity threshold and an expiry. Never use it for anything that depends on the account, the date or the user's permissions.
Set a context budget
A context budget caps what each task type may send: retrieved passages, tool output size and the history carried forward before it is summarized. Set it from evaluation results. Lower the cap until the pass rate moves, step back one notch, and keep the setting in versioned configuration.
Batch what can wait
Overnight classification, backfills and evaluation runs belong in a batch tier. The FinOps Foundation lists batch pricing among AI pricing models, at reduced rates subject to availability, for workloads that tolerate interruption.8 On capacity you run yourself, batching lets each GPU serve larger batches at higher utilization.7
Cap the output
Set a maximum output length per task type, request structured output wherever a system reads the answer, and choose the lowest reasoning effort that passes. Reasoning is billed as output, so a model thinking at full effort on a simple lookup pays for reasoning the task never needed.
| Lever | Where it fits | Quality risk | Proof it held |
|---|---|---|---|
| Routing | Mixed workloads, many easy requests | Hard cases sent to a small model | Pass rate per tier on graded cases |
| Prompt caching | Long, stable instructions; multi-step agents | Low; the model sees the same input | Cache hit rate and cost per task |
| Semantic caching | Answers identical for every user | A stale answer to a near match | Sampled review of cache hits |
| Context budget | Retrieval-heavy, long-running tasks | Dropping the passage that held the answer | Pass rate as the cap tightens |
| Batching | Work that can wait hours | Late results when the tier is busy | Completion time against the deadline |
| Output caps | Every task type | Truncated answers | Rate of truncated or invalid outputs |
Model portability as a cost lever
The ability to move a workload to a cheaper model is worth money only when a test can prove the cheaper model still does the job. Stanford's AI Index reported open-weight models narrowing the gap with closed models from 8% to 1.7% on some benchmarks,6 and in the September 2026 survey, 51% of large companies placed their model mix at the heavily frontier end while only 24% expected to be there a year later.10 Teams that can switch capture each price drop.
Portability takes three things: an interface in your code that hides which model sits behind it, prompts versioned per model, and a graded evaluation set you can run against any candidate in an afternoon. The third decides it, because a lower listed price can mean a higher bill.2 Our guide to LLM and AI agent evaluation covers how to build that set.
Self-hosting vs API: the variables that decide it
Self-hosting an open-weights model beats paying per token only when your capacity stays busy, your volume is steady and you have the people to run inference servers; for spiky or small workloads, the API is cheaper. Owned or reserved capacity costs the same idle or loaded, so its real price per token is its fixed cost divided by the tokens it serves. A FinOps Foundation working group reports GPUs often running at 15 to 30% of capacity,7 and at a quarter of capacity each token costs four times what it would on a fully used cluster. The Foundation notes that reserved throughput bought from a provider can sit underused in the same way.8
| Variable | Favors a per-token API | Favors self-hosting or reserved capacity |
|---|---|---|
| Utilization | Spiky, seasonal or low average load | High, steady load around the clock |
| Volume | Small, uncertain or still growing | Large and predictable for a year or more |
| Model need | Only a frontier model passes | An open-weights model passes your graded cases |
| Latency and location | Standard response times are fine | Strict latency, edge sites or disconnected operation |
| Data residency | Provider regions and terms meet your rules | Data may not leave your environment |
| People | No team to run GPUs and carry a pager | A platform team already runs inference |
| Commitment | You need freedom to switch models | You can commit to hardware or a capacity term |
Write the break-even down before buying hardware: the monthly fixed cost of owned capacity, including power, software, upgrades and staff time, divided by the per-token price of an API model that passes the same tests. That is the monthly token volume above which owning wins, for as long as utilization holds. We usually recommend starting on an API, metering everything and revisiting once a year of traffic shows a steady base load.
Per-agent budgets and kill switches
Every agent should run under a budget enforced in code, outside the model, with limits per call, per task, per day and per month, and a switch that stops it. An agent in a loop spends at machine speed. OWASP lists unbounded consumption, including denial of wallet attacks that exploit pay-per-use pricing, and recommends rate limits, quotas, timeouts and monitoring.3 The FinOps Foundation pairs usage limits, quotas and throttling with anomaly detection.8
Companies governing developer AI tools already rely mainly on spending caps and token limits.10 An account-level cap, though, stops every workload when it trips. Put the budget in your gateway and orchestration code, where it can stop one agent and leave the rest running.
- Per callToken limits and a timeout on every model call.On breach, the call fails fast and retries once after validation.
- Per taskLimits on steps, tool calls, tokens and elapsed time.On breach, the agent hands the case to a person, context attached.
- Per agent, dailyA daily budget, with an alert to the owner as spend nears it.On breach, non-urgent work queues until the next day.
- Per agent, monthlyA cap agreed with finance and tied to the forecast.On breach, the agent falls back to a cheaper tier or pauses; the owner approves any raise.
- Kill switchOne action, open to on-call engineering and the business owner, suspends the agent.Tested before go-live and after every major change.
Alert on cost per task and tokens per step as well as on totals: a loop shows up in tokens per step within minutes, long before it moves the daily total.
Showback and chargeback for AI spend
Showback reports each team's AI spend back to it; chargeback books that spend against the team's budget. The FinOps Foundation separates them by formality: neither is more mature than the other, showback is always required, and chargeback depends on your accounting policies.4
AI allocation is harder than cloud allocation because one invoice and a few shared API keys can cover dozens of teams. Route every model call through a gateway that stamps it with metadata, then allocate from those records:
- Team and cost center who pays.
- Product and use case which business outcome the spend supports.
- Agent and version which agent, prompt version and model made the call.
- Task type and task id so cost per completed task comes from records.
- Environment production, test and evaluation, kept apart.
- Principal the person or service account the agent acted for.
Shared costs such as the gateway, evaluation runs and the vector store need a written allocation rule, by share of tokens or of tasks, agreed with finance before the first report goes out.
Billing data is catching up. FOCUS, the FinOps Foundation's open specification for cost and usage data, is at version 1.4. Version 1.5, planned for ratification on December 3, 2026, is scoped to name the model billed, separate cached from fresh tokens, and identify who or what consumed them, down to an autonomous agent.11 Ask your providers when they will emit it. In the September 2026 survey, 23% of buyers asked providers for more transparency and granular data; 4% asked for lower prices.10
The goal is a conversation finance can hold. In that survey of 472 large companies, 39% were not confident they could connect AI spend to a measurable business outcome their CFO would accept.10 Showback by task, with cost per completed task beside value per task, is how a team leaves that group.
A monthly AI cost review
Hold the review monthly with the product owner, the engineering lead and finance, and read unit cost before total cost. The 2026 State of FinOps survey found many organizations being asked to fund AI through optimization savings,5 and this review is where those savings are found and proven.
The monthly AI cost review, in order
- Cost per completed task by system and task type, median and 95th percentile, against last month and the forecast.
- The change in total spend, split into volume, mix, tokens per task, effective price and pass rate.
- Tokens per task by category, from instructions to retries.
- Cache hit rate, and any release that broke the cache.
- Traffic share by routing tier, with each tier's pass rate.
- Every budget breach, throttle and kill-switch event, with its cause and fix.
- Unallocated spend, with an owner to fix each missing tag.
- One decision: next month's lever, its owner and the quality test it must pass.
Not every system needs every lever. If an assistant costs less each month than the engineering time it would take to tune it, meter it, cap it and leave it alone. Optimization pays where volume is high or growing, and the review shows which systems those are.
This is operating work for as long as the system runs. Our cloud and DevOps team builds the gateway, metering and budgets into the platform, in your cloud account or your data center. Where you would rather hand over the monthly review and the levers, managed operations can own them and report cost per completed task to finance every month.
Questions leaders ask
How much does it cost to run an LLM in production?
It depends on four numbers you can measure: tasks per month, tokens per task, the effective price per token after caching and discounts, and the share of tasks that pass your quality check, plus fixed platform and review costs. A price list covers only one of them. Run a pilot on real traffic, log tokens by category, and compute cost per completed task before you forecast the full rollout.
How do I reduce LLM costs without losing quality?
Measure first, then pull the lever that matches where your tokens go: route easy tasks to smaller models, cache stable prompt prefixes, cap retrieved context and output length, and batch work that can wait. Ship each change behind your evaluation set and keep it only if cost per completed task falls while the pass rate holds. A cut that raises escalations or failures is a cost increase in disguise.
How much does an AI agent cost per month to run?
Multiply the tasks it handles in a month by the tokens each task consumes across all its steps, at your effective price, then add retrieval, tool calls, platform and the human time spent on its escalations. Agents vary more than chat systems because the number of steps changes from task to task, and a 2026 study found reasoning tokens for the same query varying up to 9.7x between runs.2 Budget from a shadow run on live traffic.
Is self-hosting an LLM cheaper than using an API?
Only when the capacity stays busy. Owned or reserved capacity costs the same idle or loaded, so its real price per token rises as utilization falls, and a FinOps Foundation paper reports GPUs often running at 15 to 30% of capacity.7 Self-hosting wins with high, steady volume, an open-weights model that passes your tests, strict residency or latency needs and a team that already runs inference. Otherwise, start on an API.
What is FinOps for AI?
FinOps for AI applies the FinOps practice of visibility, allocation, optimization and unit economics to spending on models, tokens, GPUs and AI services. It is the top forward-looking priority for FinOps teams, and 98% of respondents to the 2026 State of FinOps survey now manage AI spend.5 In practice it means metering every call, allocating it to an owner and managing cost per outcome alongside the total bill.
What is the difference between showback and chargeback for AI costs?
Showback reports each team's AI spend to that team; chargeback posts the same spend to the team's official budget. The FinOps Foundation treats showback as required in every practice and chargeback as a choice that depends on your accounting policies.4 For AI, start with showback by team, product and agent, and add chargeback once the allocation rules for shared costs have held for a few cycles.
Sources
- LLM responses to benchmark questions are getting longer over timeEpoch AI, April 17, 2025
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost MoreChen, Zhang, He, Stoica, Zaharia and Zou, arXiv 2603.23971, March 2026, revised May 2026
- LLM10:2025 Unbounded ConsumptionOWASP Top 10 for LLM Applications 2025
- Invoicing & ChargebackFinOps Foundation, FinOps Framework capability
- State of FinOps 2026 ReportFinOps Foundation, February 2026
- The 2025 AI Index ReportStanford Institute for Human-Centered Artificial Intelligence, 2025
- Optimizing GenAI Usage: A FinOps Perspective on Cost, Performance, and EfficiencyFinOps Foundation, FinOps for AI Working Group, May 23, 2025
- FinOps for AI OverviewFinOps Foundation, updated February 17, 2026
- RouteLLM: Learning to Route LLMs with Preference DataOng et al., arXiv 2406.18665, June 2024
- State of Tokenomics, September 2026Tokenomics Foundation, a Linux Foundation project, September 23, 2026
- FOCUS 1.5 Release Scope: Confirmed and Stretch FeaturesFinOps Foundation, FOCUS working group, status as of September 17, 2026
- Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic TasksLumer et al., arXiv 2601.06007, January 2026
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.