LLM and RAG engineering Playbook

Running AI in Production: How to Operate AI Systems After Go-Live

An AI system starts changing the day it launches. The provider updates and retires the model under it, its users find questions nobody tested, its bill is metered and the law now expects someone to watch it. In the public record, customers and journalists found most failures first, and the operating model in this guide is built to find them sooner.

For CTOs, COOs and heads of engineering with AI systems or agents in production, deciding how to monitor and support them and who should own that work.

Published
Reviewed
Reading time
18 min

The short answer

Running AI in production means holding quality to a target as well as uptime. Pin model versions and track their retirement dates, trace every request and its cost, score sampled live traffic against quality objectives, cap spend and agent actions, and rehearse rollback and incident reporting. Give all of it named owners on call, because users, journalists or the bill found most public AI failures first.

Key takeaways

  • In the public record, operators' own monitoring caught the hard outages, while customers, reporters or the invoice surfaced wrong, harmful and degraded answers. One serving bug at Anthropic degraded answers for about a month before user reports exposed it.1
  • The model under a product retires on its provider's timetable. Anthropic gives at least 60 days' notice, and Microsoft retires generally available models 18 months after launch with no extensions.23
  • Pinning a model version leaves room for silent change: a 2026 study of 16 providers found none that let outsiders verify the served model matches its documentation.4
  • Quality needs its own objectives and its own on-call. In LangChain's survey of agent builders, 52% evaluate before release and 37% evaluate live traffic.5
  • Regulation now writes operations into law: EU deployers of high-risk AI must monitor, keep logs for at least six months and suspend use when they suspect a risk, and providers report serious incidents within 15 days.67
  • Gartner expects running a model to be at least 70% of its lifetime cost, and at least half of generative AI projects to overrun their budgets by 2028 through poor architecture and a lack of operational know-how.8

AI budgets are written for the build. The launch gets a date and a demo, and then the system meets what no pilot tests: time. In the months after go-live the provider updates the model, announces its retirement and has an outage or two, the users ask things nobody put in the test set, and the token bill grows with every new feature. The AI Incident Database recorded 362 incidents in 2025, up from 233 the year before.9 This playbook covers how AI systems fail once they are live, who notices, and the operating model that keeps them reliable: what to measure, how to manage model changes, what the rules now require and who should own the work.

  • 30 daysa serving bug degraded a frontier model's answers before user reports exposed it; the provider's own evaluations were too noisy to catch it1
  • 60 daysminimum notice Anthropic gives before retiring a model; seven of its model IDs were retired in the first nine months of 20262
  • 37%of teams building agents evaluate live traffic, against 52% who evaluate offline before release5

How AI systems fail after go-live

The public record of failures after launch falls into a few patterns. The provider changes the model or the serving stack underneath the product, or goes down. A customer-facing assistant states a policy or a law that does not exist, and the deployer is held to it. Users talk a bot into commitments nobody meant it to make. An agent with production access destroys data, or a leaked key runs up a bill. The table lists cases from each pattern, how each came to light and the control that would have caught it sooner.

IncidentWhat happenedFound byThe control that was missing
GPT-4o update, April 2025An update made the model flatter users and agree with harmful choices; OpenAI said it had focused too much on short-term feedback, and rolled the update back four days after release10Users on social mediaBehaviour regression tests as a release gate, and a staged rollout
Anthropic, August to September 2025Three infrastructure bugs degraded Claude's answers; the worst sent up to 16% of one model's requests in a single hour to the wrong servers, and it ran for about a month1User reportsContinuous quality evaluation of production traffic
OpenAI, December 2024A new telemetry service overloaded the clusters' control planes for more than four hours, and engineers were locked out of the control plane they needed to roll it back11The provider's monitoringPhased rollout, and emergency access that does not depend on the failing system
OpenAI, June 2025Elevated error rates across the API and ChatGPT for about 15.5 hours12The provider's monitoringA tested fallback to a second model or region
Air Canada, decided February 2024The website chatbot promised a retroactive bereavement fare the airline did not offer; a tribunal held the airline liable and ordered it to pay $81213The customer's claimAnswers grounded in the policy source, and a person for anything involving money
NYC MyCity, 2023 to 2024The city's business chatbot told owners they could take workers' tips and refuse housing vouchers, contrary to city law, about five months after launch14Journalists at The MarkupExpert review of sampled answers, and refusal on legal questions without a cited source
DPD, January 2024After a system update, the parcel firm's chatbot swore at a customer and criticized the company; DPD disabled the AI element15A customer's post on XGuardrail regression tests after every update (the kill switch worked)
Chevrolet dealer, December 2023A user instructed a dealer's chatbot to agree with everything and it agreed to sell a new Tahoe for $116Users, then a traffic spikeHard limits on what the bot can commit to, and alerts on unusual volume
DataTalks.Club, February 2026An AI coding agent ran a destroy command against production, deleting the course platform's database, including a table of 1,943,200 rows; the restore took about 24 hours17The operatorDeletion protection, tested restores, and no destroy rights for agents
Stolen API key, February 2026A thief used a leaked Gemini key to spend $82,314.44 in 48 hours at a company that usually spent $180 a month18The billHard spend caps per key, and alerts on spend anomalies
Vendor postmortems, a tribunal ruling, first-hand accounts and press reporting. Each case maps to a control that already exists.

Read the third column. Operators' monitoring spotted the hard outages. Almost everything else, including wrong answers, harmful answers and a model quietly getting worse, was found by customers, reporters or the invoice. Anthropic's postmortem is candid about why: its evaluations were too noisy to separate the broken serving paths from the working ones, and privacy controls limited what engineers could see of the failing conversations. It has since committed to running its evaluations continuously on production systems.1

Exhibit 1How long each failure ran before it was found or fixed
  • DPD chatbotabout a day
  • GPT-4o update4 days
  • Anthropic serving bugabout 30 days
  • NYC MyCity chatbotabout 5 months
Days from release or launch to discovery or fix, from the sources in the table above. In every case people outside the operator noticed first. Sources: [15], [10], [1], [14]

When the assistant is wrong, the deployer pays. The tribunal in the Air Canada case refused the argument that the chatbot was responsible for its own words.13 Klarna, which had moved much of its customer service to an AI assistant, said in 2025 that cost had been a too predominant factor and that the result was lower quality, and began recruiting people again.19 A system that answers customers has to be held to a quality standard as firmly as to a cost one.

The model under you changes, then retires

A production AI system usually runs on a model somebody else owns, and the owner changes it on its own schedule. Every major provider now publishes a lifecycle policy. The notice periods leave little slack for evaluating a replacement, testing it and releasing it through change control.

ProviderNotice before retirementWhat else to plan for
OpenAI APIAt least 6 months for generally available models, 3 months for specialized variants, as little as 2 weeks for previews20Aliases ending in -latest are retired too, at about three months' notice; one model announced in September 2026 had about 19 days
Anthropic APIAt least 60 days for publicly released models2Every retirement since December 2025 has had 60 to 62 days' notice; Bedrock and Google Cloud set their own dates
Microsoft FoundryAt least 60 days for generally available models, 30 for previews3Models retire 18 months after launch, 12 for some partner models; dates are not extendable, and Standard deployments upgrade automatically unless set otherwise
Amazon BedrockA Legacy period of 6 months, or 45 days for some models, before end of life21In Legacy, existing customers can lose access after 15 days without use, and extended access near the end costs more22
Provider policies as published in October 2026. Notice periods and lifecycles change; read them into a model inventory rather than a slide.

Retirement is routine. Anthropic retired seven model IDs in the first nine months of 2026.2 The same model can also retire on different dates in different places: Claude Sonnet 4 left Anthropic's own API on June 15, 2026, and stays on Amazon Bedrock until October 14, 2026.222 Microsoft tells customers to start evaluating newer models against their own prompts and data without waiting for an official replacement, which it names only about 90 to 120 days before retirement.3 The working assumption for planning is a life of roughly 12 to 18 months per model version, with at least one forced migration in the life of any system.

Pinning a dated version protects against only part of the change. In 2023, Stanford and Berkeley researchers found GPT-4's accuracy at identifying prime numbers fell from 84% to 51% between March and June, and the share of its code that ran as returned fell from 52% to 10%; they named weaker instruction-following as a common factor behind the drift, and instruction-following is what an automated pipeline depends on most.23 Anthropic's 2025 bugs changed answers with no change of model version at all.1 An August 2026 study of 16 providers found none that let an outside party verify that the model being served is the one its documentation describes.4 The defence is a fixed set of your own cases run against production on a schedule, with an alert on any shift, and every replacement model evaluated on the day it ships.

Availability needs the same realism. Amazon Bedrock's agreement commits to 99.9% monthly uptime in each region, with service credits of 10% to 100% of the month's bill, and excludes downtime caused by an inoperable model.24 A 99.9% month still allows about 43 minutes of downtime, and OpenAI's June 2025 incident lasted about 15.5 hours.12 A credit is a share of the model bill, which is usually small next to what an outage costs the business, so production systems need a tested fallback to a second model or region.

Hold quality to an objective, the way uptime is

An AI service can answer every request quickly and cheaply with zero errors and still be wrong. Grafana Labs makes that point and proposes the fix site reliability teams will recognize: define an indicator for each behaviour that matters, such as grounded answers or completed tasks, score it with an evaluator on sampled conversations, and give each its own objective, for example 95% fulfilled over 30 days.25 Google's SRE workbook already defines an indicator as good events over total events, allows correctness as the good event, and includes error-budget policies that freeze changes until a service is back within its objective.26

The cloud providers' guidance converges on the same layers. AWS rates missing foundation-model monitoring a high risk and lists invocations, latency, token usage, errors and throttling, with incident playbooks practised for when the alarms fire.27 Microsoft adds continuous evaluation of production traffic at a sampled rate, scheduled re-runs of test sets to detect drift and scheduled red teaming.28 Practice lags the guidance: of the more than 1,300 practitioners LangChain surveyed, 52.4% run offline evaluations and 37.3% evaluate live traffic.5 Our guide to LLM evaluation covers building the test set and the release gate.

Datadog's data from its customers shows where everyday failures sit. In February 2026, 5% of model calls returned an error and 60% of those errors were rate limits, so capacity handling, backoff and fallback routing come first.29 The same data shows system prompts making up 69% of input tokens while only 28% of calls on models that support caching read any cached input, which is spend that tuning can recover.29 Datadog sells observability, and these figures come from its own customer base.

Instrument with OpenTelemetry so traces outlive any one vendor's dashboard. Its conventions for model calls, agents and tool calls are still in development, and in June 2026 they moved out of the core specification into a repository of their own.30 Pin the version you emit and keep a collector that can remap names when they change.

ObjectiveIndicatorWhat it catches
AvailabilityShare of requests that succeed, with rate-limited requests counted as failuresProvider outages and capacity limits
LatencyTime to first token and end-to-end time at the 95th percentile, per taskSlow providers and long agent chains
QualityShare of sampled responses passing each evaluator: grounded, task completed, correct formatDrift, bad updates, flattering or wrong answers
SafetyShare passing policy checks, with zero tolerance for leaked secrets or unauthorized commitmentsManipulated bots and harmful output
CostCost per completed task against its budgetAgent loops, leaked keys, swollen prompts
LifecycleDays until each model in production retires, and whether its replacement has passed evaluationForced migrations under deadline
A quality indicator inherits the error of the judge that scores it. Calibrate automated judges against human labels before anyone is paged on them.

Limits that hold at machine speed

Agents and keys fail faster than people react. The $82,314.44 Gemini bill accrued over two days.18 The DataTalks.Club database went in one command, and its founder wrote afterwards that delegating the plan, apply and destroy steps had removed the last safety layer. He then enabled deletion protection, daily automated restore tests and backups kept apart from the infrastructure tooling.17

The controls are ordinary, and they have to be hard limits rather than dashboards: spend caps per key and per task, least-privilege credentials, approval before any irreversible action, and an iteration limit on every agent loop, a stopping condition Anthropic's own agent guidance recommends.31 The switch that turns a feature off must work without the system it stops. In OpenAI's December 2024 outage, the fix depended on a control plane the failure had made unreachable.11 Our AI agent guardrails playbook covers these controls in depth.

What the rules now expect after launch

Regulation has caught up with operations. The EU AI Act requires deployers of high-risk systems to use them as instructed, assign oversight to people with the competence and authority to act, monitor the system, suspend use when they suspect a risk and keep its logs for at least six months.6 Providers must run a documented post-market monitoring system and report serious incidents, and they may not alter the system in a way that could affect the investigation before telling the authorities, which an incident runbook has to allow for.327 The Digital Omnibus, in force since July 27, 2026, moved the high-risk dates to December 2, 2027 for Annex III systems and August 2, 2028 for Annex I, and left the incident deadlines as they were.33

RuleWhat it asks after launchClock
EU AI Act, deployers of high-risk systemsCompetent human oversight, monitoring, suspension when a risk is suspected, logs kept at least six months6Annex III systems from December 2, 202733
EU AI Act, providersPost-market monitoring, serious-incident reports, no alteration before the authorities are told32715 days; 2 days for widespread or critical-infrastructure incidents; 10 days after a death
NIST AI RMF, MANAGE 4.1Post-deployment monitoring plans covering user input, appeal and override, decommissioning, incident response, recovery and change management34Voluntary
DORA, EU financial entitiesMajor ICT incident reports, which cover the ICT services an AI provider supplies35Initial notice within 4 hours of classification and 24 hours of awareness; then 72 hours; then one month
US bank model risk, SR 26-2Replaced SR 11-7 on April 17, 2026, and leaves generative and agentic AI outside its scope for now3637A request for information on AI is promised
Colorado SB 26-189Notice before automated decisions, an explanation after an adverse one, human review on request, records kept three years38From January 1, 2027
Post-deployment duties as of October 2026. The EU dates follow the Digital Omnibus; check sector rules where you operate.

The direction is the same in every regime: an inventory, logs, monitoring, a person who can override, an incident path and accountability for vendor models. A provider's retirement schedule becomes the deployer's compliance problem, and a model swap becomes a change to assess as well as to test. Our EU AI Act guide covers classification and the obligations in detail.

Running costs more than building

Gartner expects inference to account for at least 70% of a model's lifetime cost, warns that a production-ready system can cost orders of magnitude more than the pilot, and predicts that at least half of generative AI projects will overrun their budgets by 2028 through poor architectural choices and a lack of operational know-how.8 None of that is new to machine learning. Google's engineers wrote in 2015 that real-world ML systems commonly incur massive ongoing maintenance costs, and that the model code is a small part of the system around it.39 Our guides to inference cost and moving a pilot to production cover the economics in more depth.

The skills are scarce. Deloitte's 2026 survey of 3,235 leaders found talent the least prepared area for AI, with 20% calling their organization highly prepared, and one AI leader it interviewed discovered there was no clear inventory of the models running in production.40 ManpowerGroup's survey of 39,063 employers ranks AI model and application development as the hardest skill to hire.41 Gartner puts 2026 spending on AI services at $576.5 billion, ahead of AI software at $461.6 billion.42 Incident response is maturing too: OWASP published a generative AI incident response guide in July 2025 and the Coalition for Secure AI a framework in November 2025, both written for security teams.4344 Quality incidents, such as a slow slide in answer accuracy, still need product-level owners and runbooks.

Exhibit 2Three ways to own AI after launch

Each product team

Whoever built it runs it

  • Close to the users and the domain
  • AI on-call added to people hired to build features
  • Every team relearns evaluation, lifecycle and cost
  • No single inventory of what is running

Central platform team

Shared tooling, shared on-call

  • One gateway, tracing, evaluation and model inventory
  • Scarce skills concentrated in one place
  • Can become a queue between teams and production
  • Product owners still set the quality targets

Managed service

What we run

  • Named engineers on call against agreed objectives
  • Model lifecycle, evaluations and cost reviewed monthly
  • The client keeps the decisions and the access
  • Runbooks and documentation handed over in full
Under every model, accountability for a deployed system stays with the organization that deploys it. An operator runs the work; the deployer answers for it.

An operating model that finds problems first

The practices that close the gaps in the incident table are known and unglamorous. They work in this order, because each one depends on the one before it.

  1. Inventory what is running Every model and version, prompt, tool, agent and key in production, each with an owner and the retirement date of its model, read from the providers' lifecycle data rather than from memory.
  2. Trace every request OpenTelemetry traces across model calls, tool calls and agent steps, with tokens and cost per task, pinned to a known convention version and routed through a collector you control.
  3. Set objectives for quality as well as uptime A test set built from your own cases, sampled evaluation of live traffic, an objective per behaviour, and an error-budget policy that freezes prompt and model changes when quality falls below target.
  4. Put limits where machines outrun people Spend caps per key and per task, iteration limits on agents, least-privilege credentials, approval for irreversible actions, and a kill switch that works when the system it stops is down.
  5. Rehearse the bad day Runbooks for a provider outage, a quality regression, a harmful answer and a cost spike, each with its fallback or rollback tested, and an incident path that meets the regulatory clocks you are under.
  6. Treat every model change as a release Evaluate candidate models the day they ship, roll out by canary with a pinned rollback, and repeat the risk assessment before the provider's retirement date forces the move.

Questions to ask whoever runs your AI

  • Which model versions are in production today, and when does each one retire?
  • Who is paged when answer quality drops, and which number pages them?
  • How would we know if the provider changed the model's behaviour without changing its name?
  • What stops an agent or a leaked key from spending a month's budget in a day?
  • Can we switch off one AI feature in minutes without taking the product down?
  • If a serious incident happened tonight, who files the report, and within which deadline?
  • When did we last restore production data from backup, and how long did it take?

Launch is the point where an AI system starts depending on things its builders do not control: a provider's release and retirement calendar, the questions real users ask, a metered bill and a regulator's clock. The organizations that keep their AI reliable treat it like any production service with a few extra dials, and they decide before launch who watches those dials at three in the morning.

This is the work of our managed services. Our engineers run AI systems in production against agreed objectives for quality, latency and cost: every model version tracked to its retirement date, live traffic evaluated, spend and agent actions capped, on-call with tested runbooks, and a monthly report that puts quality beside uptime.

Questions leaders ask

What does it mean to run AI in production?

It means operating an AI system after launch the way you operate any production service, with a few additions: tracking the provider's model versions and retirement dates, evaluating sampled live traffic for quality, capping spend and agent actions, tracing every request with its cost, and keeping on-call engineers with tested runbooks for outages, quality regressions and incidents.

How do you monitor an LLM application in production?

Trace every request with OpenTelemetry, recording model, tokens, latency, tool calls and cost per task. Track availability, latency and rate limits, and add quality indicators: the share of sampled responses that pass evaluators for groundedness, task completion, format and safety. Re-run a fixed test set on a schedule to catch drift, and alert when any indicator falls below its objective.

What happens when an AI provider retires a model?

Requests to the retired model fail. Anthropic gives at least 60 days' notice, OpenAI at least six months for generally available models and much less for previews, and Microsoft retires generally available models 18 months after launch with no extensions. Keep an inventory of models and dates, evaluate replacements as soon as they ship and migrate by canary with a rollback.

What SLOs should an AI system have?

Six families cover most systems: availability with rate limits counted as failures, latency including time to first token, quality as the pass rate of sampled responses against each evaluator, safety with zero tolerance for leaked secrets, cost per completed task, and lifecycle, meaning days until each model retires. Calibrate automated judges against human labels before paging on them.

Who should own AI systems after launch?

Someone named, with on-call duty and authority to roll back. Product teams know the users but tend to lack AI operations skills; a central platform team concentrates tooling and skills but can become a queue; a managed service brings a run team against agreed objectives. In every case the deploying organization keeps the accountability.

What does the EU AI Act require after an AI system is deployed?

For high-risk systems, deployers must use them as instructed, assign competent human oversight, monitor operation, suspend use when they suspect a risk and keep logs for at least six months. Providers run post-market monitoring and report serious incidents within 15 days, 2 days for widespread incidents and 10 days after a death. The Digital Omnibus moved Annex III obligations to December 2, 2027.

How do you stop AI agents and API keys from running up costs?

Use hard limits: spend caps per key and per task, iteration limits on every agent loop, rate limits, and keys restricted to the services they need. Alert on spend anomalies within hours, since a stolen key ran up $82,314.44 in 48 hours. Report cost per completed task beside quality, so savings never come from quietly worse answers.

Sources

  1. A postmortem of three recent issuesAnthropic Engineering, September 2025
  2. Model deprecationsAnthropic, Claude Developer Platform documentation
  3. Foundry Models lifecycle and support policyMicrosoft Learn, July 2026
  4. Silent Updates: Measuring and Closing the Post-Deployment Disclosure GapAbraham and Bucknall, arXiv 2608.11803, August 2026
  5. State of Agent EngineeringLangChain, survey of more than 1,300 professionals, December 2025
  6. Article 26: Obligations of Deployers of High-Risk AI SystemsEU Artificial Intelligence Act
  7. Article 73: Reporting of Serious IncidentsEU Artificial Intelligence Act
  8. Gartner: Half of Gen AI Projects Could Exceed Budget by 2028Campus Technology, June 22, 2026
  9. The 2026 AI Index Report: Responsible AIStanford HAI, 2026
  10. OpenAI pulls plug on overly supportive ChatGPT smarmbotThe Register, April 30, 2025
  11. API, ChatGPT and Sora facing issuesOpenAI status incident report, December 11, 2024
  12. Elevated error ratesOpenAI status incident, June 10, 2025
  13. Air Canada found liable for chatbot's bad advice on plane ticketsCBC News, February 2024
  14. NYC's AI Chatbot Tells Businesses to Break the LawThe Markup, March 29, 2024
  15. DPD chatbot goes off the rails at suggestion of customerThe Register, January 23, 2024
  16. Incident 622: Chevrolet Dealer Chatbot Agrees to Sell Tahoe for $1AI Incident Database
  17. How I Dropped Our Production Database and Now Pay 10% More for AWSAlexey Grigorev, DataTalks.Club, 2026
  18. Dev stunned by $82K Gemini API key bill after theftThe Register, March 3, 2026
  19. Klarna changes its AI tune and again recruits humans for customer serviceCX Dive, May 2025
  20. DeprecationsOpenAI API documentation
  21. Model lifecycleAmazon Bedrock User Guide
  22. Model lifecycle (models launched before September 7, 2026)Amazon Bedrock User Guide
  23. How Is ChatGPT's Behavior Changing over Time?Chen, Zaharia and Zou, Stanford and UC Berkeley, arXiv 2307.09009
  24. Amazon Bedrock Service Level AgreementAmazon Web Services
  25. What if your agent's hallucinations had a budget? How to start using SLOs for agent behaviorGrafana Labs, 2026
  26. Implementing SLOsGoogle, The Site Reliability Workbook
  27. GENOPS02-BP02 Monitor foundation model metricsAWS Well-Architected Generative AI Lens
  28. Observability in generative AIMicrosoft Learn, Microsoft Foundry, July 2026
  29. State of AI EngineeringDatadog, July 2026
  30. Semantic Conventions v1.42.0 release notesOpenTelemetry, June 2026
  31. Building effective agentsAnthropic Engineering, December 19, 2024
  32. Article 72: Post-Market Monitoring by Providers and Post-Market Monitoring Plan for High-Risk AI SystemsEU Artificial Intelligence Act
  33. Digital Omnibus on AIEU Artificial Intelligence Act explorer, 2026
  34. AI RMF Playbook: ManageNational Institute of Standards and Technology
  35. Commission Delegated Regulation (EU) 2025/301 on the content and time limits for reporting major ICT-related incidentsEUR-Lex, 2025
  36. Supervisory Letter SR 26-2 on Revised Guidance on Model Risk ManagementBoard of Governors of the Federal Reserve System, April 17, 2026
  37. Visual memo: Key changes under the federal banking agencies' revised model risk management guidanceDavis Polk, 2026
  38. Colorado Enacts New Law Regulating Automated Decision-Making TechnologyLathrop GPM, June 1, 2026
  39. Hidden Technical Debt in Machine Learning SystemsSculley et al., Google, NeurIPS 2015
  40. The State of AI in the Enterprise, 2026Deloitte, survey of 3,235 leaders
  41. Global Talent Shortage Reaches Turning Point as AI Skills Claim Top SpotManpowerGroup, February 26, 2026
  42. Gartner: Worldwide AI Spending to Reach $2.67 Trillion in 2026THE Journal, September 21, 2026
  43. GenAI Incident Response Guide 1.0OWASP GenAI Security Project, July 28, 2025
  44. Coalition for Secure AI Releases Two Actionable Frameworks for AI Model Signing and Incident ResponseOASIS Open, November 18, 2025

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 4, 2026. No client data appears in our insights.

Read next

All insights
  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • On one plinth, a small glass prototype sits on a workbench at the front and drums of live data wait behind it; both run along lit conduits into a glowing go-live gate under a shield, and once it passes them a lit production tower rises at full height under a beam of light.

    Economics and buying Playbook

    AI Pilot to Production: How to Rescue, Restart, Buy or Stop a Stalled Pilot

    For CTOs, CIOs and heads of AI with a pilot that worked in the demo and has not been signed off to run for real, deciding whether to rescue it, rebuild it, buy instead or stop.

    22 min read

  • Three AI agents send their calls through one lit router, which passes most of them to a fleet of small models and only a hard one to a large frontier model. Each agent has its own budget gauge; one has spent to its cap and a red barrier stops it, while the other two keep working.

    Economics and buying Playbook

    AI Inference Cost: How to Govern LLM and Agent Spend

    For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.