LLM and RAG engineering Playbook
Running AI in Production: How to Operate AI Systems After Go-Live
An AI system starts changing the day it launches. The provider updates and retires the model under it, its users find questions nobody tested, its bill is metered and the law now expects someone to watch it. In the public record, customers and journalists found most failures first, and the operating model in this guide is built to find them sooner.
For CTOs, COOs and heads of engineering with AI systems or agents in production, deciding how to monitor and support them and who should own that work.
The short answer
Running AI in production means holding quality to a target as well as uptime. Pin model versions and track their retirement dates, trace every request and its cost, score sampled live traffic against quality objectives, cap spend and agent actions, and rehearse rollback and incident reporting. Give all of it named owners on call, because users, journalists or the bill found most public AI failures first.
Key takeaways
- In the public record, operators' own monitoring caught the hard outages, while customers, reporters or the invoice surfaced wrong, harmful and degraded answers. One serving bug at Anthropic degraded answers for about a month before user reports exposed it.1
- The model under a product retires on its provider's timetable. Anthropic gives at least 60 days' notice, and Microsoft retires generally available models 18 months after launch with no extensions.23
- Pinning a model version leaves room for silent change: a 2026 study of 16 providers found none that let outsiders verify the served model matches its documentation.4
- Quality needs its own objectives and its own on-call. In LangChain's survey of agent builders, 52% evaluate before release and 37% evaluate live traffic.5
- Regulation now writes operations into law: EU deployers of high-risk AI must monitor, keep logs for at least six months and suspend use when they suspect a risk, and providers report serious incidents within 15 days.67
- Gartner expects running a model to be at least 70% of its lifetime cost, and at least half of generative AI projects to overrun their budgets by 2028 through poor architecture and a lack of operational know-how.8
AI budgets are written for the build. The launch gets a date and a demo, and then the system meets what no pilot tests: time. In the months after go-live the provider updates the model, announces its retirement and has an outage or two, the users ask things nobody put in the test set, and the token bill grows with every new feature. The AI Incident Database recorded 362 incidents in 2025, up from 233 the year before.9 This playbook covers how AI systems fail once they are live, who notices, and the operating model that keeps them reliable: what to measure, how to manage model changes, what the rules now require and who should own the work.
- 30 daysa serving bug degraded a frontier model's answers before user reports exposed it; the provider's own evaluations were too noisy to catch it1
- 60 daysminimum notice Anthropic gives before retiring a model; seven of its model IDs were retired in the first nine months of 20262
- 37%of teams building agents evaluate live traffic, against 52% who evaluate offline before release5
How AI systems fail after go-live
The public record of failures after launch falls into a few patterns. The provider changes the model or the serving stack underneath the product, or goes down. A customer-facing assistant states a policy or a law that does not exist, and the deployer is held to it. Users talk a bot into commitments nobody meant it to make. An agent with production access destroys data, or a leaked key runs up a bill. The table lists cases from each pattern, how each came to light and the control that would have caught it sooner.
| Incident | What happened | Found by | The control that was missing |
|---|---|---|---|
| GPT-4o update, April 2025 | An update made the model flatter users and agree with harmful choices; OpenAI said it had focused too much on short-term feedback, and rolled the update back four days after release10 | Users on social media | Behaviour regression tests as a release gate, and a staged rollout |
| Anthropic, August to September 2025 | Three infrastructure bugs degraded Claude's answers; the worst sent up to 16% of one model's requests in a single hour to the wrong servers, and it ran for about a month1 | User reports | Continuous quality evaluation of production traffic |
| OpenAI, December 2024 | A new telemetry service overloaded the clusters' control planes for more than four hours, and engineers were locked out of the control plane they needed to roll it back11 | The provider's monitoring | Phased rollout, and emergency access that does not depend on the failing system |
| OpenAI, June 2025 | Elevated error rates across the API and ChatGPT for about 15.5 hours12 | The provider's monitoring | A tested fallback to a second model or region |
| Air Canada, decided February 2024 | The website chatbot promised a retroactive bereavement fare the airline did not offer; a tribunal held the airline liable and ordered it to pay $81213 | The customer's claim | Answers grounded in the policy source, and a person for anything involving money |
| NYC MyCity, 2023 to 2024 | The city's business chatbot told owners they could take workers' tips and refuse housing vouchers, contrary to city law, about five months after launch14 | Journalists at The Markup | Expert review of sampled answers, and refusal on legal questions without a cited source |
| DPD, January 2024 | After a system update, the parcel firm's chatbot swore at a customer and criticized the company; DPD disabled the AI element15 | A customer's post on X | Guardrail regression tests after every update (the kill switch worked) |
| Chevrolet dealer, December 2023 | A user instructed a dealer's chatbot to agree with everything and it agreed to sell a new Tahoe for $116 | Users, then a traffic spike | Hard limits on what the bot can commit to, and alerts on unusual volume |
| DataTalks.Club, February 2026 | An AI coding agent ran a destroy command against production, deleting the course platform's database, including a table of 1,943,200 rows; the restore took about 24 hours17 | The operator | Deletion protection, tested restores, and no destroy rights for agents |
| Stolen API key, February 2026 | A thief used a leaked Gemini key to spend $82,314.44 in 48 hours at a company that usually spent $180 a month18 | The bill | Hard spend caps per key, and alerts on spend anomalies |
Read the third column. Operators' monitoring spotted the hard outages. Almost everything else, including wrong answers, harmful answers and a model quietly getting worse, was found by customers, reporters or the invoice. Anthropic's postmortem is candid about why: its evaluations were too noisy to separate the broken serving paths from the working ones, and privacy controls limited what engineers could see of the failing conversations. It has since committed to running its evaluations continuously on production systems.1
When the assistant is wrong, the deployer pays. The tribunal in the Air Canada case refused the argument that the chatbot was responsible for its own words.13 Klarna, which had moved much of its customer service to an AI assistant, said in 2025 that cost had been a too predominant factor and that the result was lower quality, and began recruiting people again.19 A system that answers customers has to be held to a quality standard as firmly as to a cost one.
The model under you changes, then retires
A production AI system usually runs on a model somebody else owns, and the owner changes it on its own schedule. Every major provider now publishes a lifecycle policy. The notice periods leave little slack for evaluating a replacement, testing it and releasing it through change control.
| Provider | Notice before retirement | What else to plan for |
|---|---|---|
| OpenAI API | At least 6 months for generally available models, 3 months for specialized variants, as little as 2 weeks for previews20 | Aliases ending in -latest are retired too, at about three months' notice; one model announced in September 2026 had about 19 days |
| Anthropic API | At least 60 days for publicly released models2 | Every retirement since December 2025 has had 60 to 62 days' notice; Bedrock and Google Cloud set their own dates |
| Microsoft Foundry | At least 60 days for generally available models, 30 for previews3 | Models retire 18 months after launch, 12 for some partner models; dates are not extendable, and Standard deployments upgrade automatically unless set otherwise |
| Amazon Bedrock | A Legacy period of 6 months, or 45 days for some models, before end of life21 | In Legacy, existing customers can lose access after 15 days without use, and extended access near the end costs more22 |
Retirement is routine. Anthropic retired seven model IDs in the first nine months of 2026.2 The same model can also retire on different dates in different places: Claude Sonnet 4 left Anthropic's own API on June 15, 2026, and stays on Amazon Bedrock until October 14, 2026.222 Microsoft tells customers to start evaluating newer models against their own prompts and data without waiting for an official replacement, which it names only about 90 to 120 days before retirement.3 The working assumption for planning is a life of roughly 12 to 18 months per model version, with at least one forced migration in the life of any system.
Pinning a dated version protects against only part of the change. In 2023, Stanford and Berkeley researchers found GPT-4's accuracy at identifying prime numbers fell from 84% to 51% between March and June, and the share of its code that ran as returned fell from 52% to 10%; they named weaker instruction-following as a common factor behind the drift, and instruction-following is what an automated pipeline depends on most.23 Anthropic's 2025 bugs changed answers with no change of model version at all.1 An August 2026 study of 16 providers found none that let an outside party verify that the model being served is the one its documentation describes.4 The defence is a fixed set of your own cases run against production on a schedule, with an alert on any shift, and every replacement model evaluated on the day it ships.
Availability needs the same realism. Amazon Bedrock's agreement commits to 99.9% monthly uptime in each region, with service credits of 10% to 100% of the month's bill, and excludes downtime caused by an inoperable model.24 A 99.9% month still allows about 43 minutes of downtime, and OpenAI's June 2025 incident lasted about 15.5 hours.12 A credit is a share of the model bill, which is usually small next to what an outage costs the business, so production systems need a tested fallback to a second model or region.
Hold quality to an objective, the way uptime is
An AI service can answer every request quickly and cheaply with zero errors and still be wrong. Grafana Labs makes that point and proposes the fix site reliability teams will recognize: define an indicator for each behaviour that matters, such as grounded answers or completed tasks, score it with an evaluator on sampled conversations, and give each its own objective, for example 95% fulfilled over 30 days.25 Google's SRE workbook already defines an indicator as good events over total events, allows correctness as the good event, and includes error-budget policies that freeze changes until a service is back within its objective.26
The cloud providers' guidance converges on the same layers. AWS rates missing foundation-model monitoring a high risk and lists invocations, latency, token usage, errors and throttling, with incident playbooks practised for when the alarms fire.27 Microsoft adds continuous evaluation of production traffic at a sampled rate, scheduled re-runs of test sets to detect drift and scheduled red teaming.28 Practice lags the guidance: of the more than 1,300 practitioners LangChain surveyed, 52.4% run offline evaluations and 37.3% evaluate live traffic.5 Our guide to LLM evaluation covers building the test set and the release gate.
Datadog's data from its customers shows where everyday failures sit. In February 2026, 5% of model calls returned an error and 60% of those errors were rate limits, so capacity handling, backoff and fallback routing come first.29 The same data shows system prompts making up 69% of input tokens while only 28% of calls on models that support caching read any cached input, which is spend that tuning can recover.29 Datadog sells observability, and these figures come from its own customer base.
Instrument with OpenTelemetry so traces outlive any one vendor's dashboard. Its conventions for model calls, agents and tool calls are still in development, and in June 2026 they moved out of the core specification into a repository of their own.30 Pin the version you emit and keep a collector that can remap names when they change.
| Objective | Indicator | What it catches |
|---|---|---|
| Availability | Share of requests that succeed, with rate-limited requests counted as failures | Provider outages and capacity limits |
| Latency | Time to first token and end-to-end time at the 95th percentile, per task | Slow providers and long agent chains |
| Quality | Share of sampled responses passing each evaluator: grounded, task completed, correct format | Drift, bad updates, flattering or wrong answers |
| Safety | Share passing policy checks, with zero tolerance for leaked secrets or unauthorized commitments | Manipulated bots and harmful output |
| Cost | Cost per completed task against its budget | Agent loops, leaked keys, swollen prompts |
| Lifecycle | Days until each model in production retires, and whether its replacement has passed evaluation | Forced migrations under deadline |
Limits that hold at machine speed
Agents and keys fail faster than people react. The $82,314.44 Gemini bill accrued over two days.18 The DataTalks.Club database went in one command, and its founder wrote afterwards that delegating the plan, apply and destroy steps had removed the last safety layer. He then enabled deletion protection, daily automated restore tests and backups kept apart from the infrastructure tooling.17
The controls are ordinary, and they have to be hard limits rather than dashboards: spend caps per key and per task, least-privilege credentials, approval before any irreversible action, and an iteration limit on every agent loop, a stopping condition Anthropic's own agent guidance recommends.31 The switch that turns a feature off must work without the system it stops. In OpenAI's December 2024 outage, the fix depended on a control plane the failure had made unreachable.11 Our AI agent guardrails playbook covers these controls in depth.
What the rules now expect after launch
Regulation has caught up with operations. The EU AI Act requires deployers of high-risk systems to use them as instructed, assign oversight to people with the competence and authority to act, monitor the system, suspend use when they suspect a risk and keep its logs for at least six months.6 Providers must run a documented post-market monitoring system and report serious incidents, and they may not alter the system in a way that could affect the investigation before telling the authorities, which an incident runbook has to allow for.327 The Digital Omnibus, in force since July 27, 2026, moved the high-risk dates to December 2, 2027 for Annex III systems and August 2, 2028 for Annex I, and left the incident deadlines as they were.33
| Rule | What it asks after launch | Clock |
|---|---|---|
| EU AI Act, deployers of high-risk systems | Competent human oversight, monitoring, suspension when a risk is suspected, logs kept at least six months6 | Annex III systems from December 2, 202733 |
| EU AI Act, providers | Post-market monitoring, serious-incident reports, no alteration before the authorities are told327 | 15 days; 2 days for widespread or critical-infrastructure incidents; 10 days after a death |
| NIST AI RMF, MANAGE 4.1 | Post-deployment monitoring plans covering user input, appeal and override, decommissioning, incident response, recovery and change management34 | Voluntary |
| DORA, EU financial entities | Major ICT incident reports, which cover the ICT services an AI provider supplies35 | Initial notice within 4 hours of classification and 24 hours of awareness; then 72 hours; then one month |
| US bank model risk, SR 26-2 | Replaced SR 11-7 on April 17, 2026, and leaves generative and agentic AI outside its scope for now3637 | A request for information on AI is promised |
| Colorado SB 26-189 | Notice before automated decisions, an explanation after an adverse one, human review on request, records kept three years38 | From January 1, 2027 |
The direction is the same in every regime: an inventory, logs, monitoring, a person who can override, an incident path and accountability for vendor models. A provider's retirement schedule becomes the deployer's compliance problem, and a model swap becomes a change to assess as well as to test. Our EU AI Act guide covers classification and the obligations in detail.
Running costs more than building
Gartner expects inference to account for at least 70% of a model's lifetime cost, warns that a production-ready system can cost orders of magnitude more than the pilot, and predicts that at least half of generative AI projects will overrun their budgets by 2028 through poor architectural choices and a lack of operational know-how.8 None of that is new to machine learning. Google's engineers wrote in 2015 that real-world ML systems commonly incur massive ongoing maintenance costs, and that the model code is a small part of the system around it.39 Our guides to inference cost and moving a pilot to production cover the economics in more depth.
The skills are scarce. Deloitte's 2026 survey of 3,235 leaders found talent the least prepared area for AI, with 20% calling their organization highly prepared, and one AI leader it interviewed discovered there was no clear inventory of the models running in production.40 ManpowerGroup's survey of 39,063 employers ranks AI model and application development as the hardest skill to hire.41 Gartner puts 2026 spending on AI services at $576.5 billion, ahead of AI software at $461.6 billion.42 Incident response is maturing too: OWASP published a generative AI incident response guide in July 2025 and the Coalition for Secure AI a framework in November 2025, both written for security teams.4344 Quality incidents, such as a slow slide in answer accuracy, still need product-level owners and runbooks.
Each product team
Whoever built it runs it
- Close to the users and the domain
- AI on-call added to people hired to build features
- Every team relearns evaluation, lifecycle and cost
- No single inventory of what is running
Central platform team
Shared tooling, shared on-call
- One gateway, tracing, evaluation and model inventory
- Scarce skills concentrated in one place
- Can become a queue between teams and production
- Product owners still set the quality targets
Managed service
What we run
- Named engineers on call against agreed objectives
- Model lifecycle, evaluations and cost reviewed monthly
- The client keeps the decisions and the access
- Runbooks and documentation handed over in full
An operating model that finds problems first
The practices that close the gaps in the incident table are known and unglamorous. They work in this order, because each one depends on the one before it.
- Inventory what is running Every model and version, prompt, tool, agent and key in production, each with an owner and the retirement date of its model, read from the providers' lifecycle data rather than from memory.
- Trace every request OpenTelemetry traces across model calls, tool calls and agent steps, with tokens and cost per task, pinned to a known convention version and routed through a collector you control.
- Set objectives for quality as well as uptime A test set built from your own cases, sampled evaluation of live traffic, an objective per behaviour, and an error-budget policy that freezes prompt and model changes when quality falls below target.
- Put limits where machines outrun people Spend caps per key and per task, iteration limits on agents, least-privilege credentials, approval for irreversible actions, and a kill switch that works when the system it stops is down.
- Rehearse the bad day Runbooks for a provider outage, a quality regression, a harmful answer and a cost spike, each with its fallback or rollback tested, and an incident path that meets the regulatory clocks you are under.
- Treat every model change as a release Evaluate candidate models the day they ship, roll out by canary with a pinned rollback, and repeat the risk assessment before the provider's retirement date forces the move.
Questions to ask whoever runs your AI
- Which model versions are in production today, and when does each one retire?
- Who is paged when answer quality drops, and which number pages them?
- How would we know if the provider changed the model's behaviour without changing its name?
- What stops an agent or a leaked key from spending a month's budget in a day?
- Can we switch off one AI feature in minutes without taking the product down?
- If a serious incident happened tonight, who files the report, and within which deadline?
- When did we last restore production data from backup, and how long did it take?
Launch is the point where an AI system starts depending on things its builders do not control: a provider's release and retirement calendar, the questions real users ask, a metered bill and a regulator's clock. The organizations that keep their AI reliable treat it like any production service with a few extra dials, and they decide before launch who watches those dials at three in the morning.
This is the work of our managed services. Our engineers run AI systems in production against agreed objectives for quality, latency and cost: every model version tracked to its retirement date, live traffic evaluated, spend and agent actions capped, on-call with tested runbooks, and a monthly report that puts quality beside uptime.
Questions leaders ask
What does it mean to run AI in production?
It means operating an AI system after launch the way you operate any production service, with a few additions: tracking the provider's model versions and retirement dates, evaluating sampled live traffic for quality, capping spend and agent actions, tracing every request with its cost, and keeping on-call engineers with tested runbooks for outages, quality regressions and incidents.
How do you monitor an LLM application in production?
Trace every request with OpenTelemetry, recording model, tokens, latency, tool calls and cost per task. Track availability, latency and rate limits, and add quality indicators: the share of sampled responses that pass evaluators for groundedness, task completion, format and safety. Re-run a fixed test set on a schedule to catch drift, and alert when any indicator falls below its objective.
What happens when an AI provider retires a model?
Requests to the retired model fail. Anthropic gives at least 60 days' notice, OpenAI at least six months for generally available models and much less for previews, and Microsoft retires generally available models 18 months after launch with no extensions. Keep an inventory of models and dates, evaluate replacements as soon as they ship and migrate by canary with a rollback.
What SLOs should an AI system have?
Six families cover most systems: availability with rate limits counted as failures, latency including time to first token, quality as the pass rate of sampled responses against each evaluator, safety with zero tolerance for leaked secrets, cost per completed task, and lifecycle, meaning days until each model retires. Calibrate automated judges against human labels before paging on them.
Who should own AI systems after launch?
Someone named, with on-call duty and authority to roll back. Product teams know the users but tend to lack AI operations skills; a central platform team concentrates tooling and skills but can become a queue; a managed service brings a run team against agreed objectives. In every case the deploying organization keeps the accountability.
What does the EU AI Act require after an AI system is deployed?
For high-risk systems, deployers must use them as instructed, assign competent human oversight, monitor operation, suspend use when they suspect a risk and keep logs for at least six months. Providers run post-market monitoring and report serious incidents within 15 days, 2 days for widespread incidents and 10 days after a death. The Digital Omnibus moved Annex III obligations to December 2, 2027.
How do you stop AI agents and API keys from running up costs?
Use hard limits: spend caps per key and per task, iteration limits on every agent loop, rate limits, and keys restricted to the services they need. Alert on spend anomalies within hours, since a stolen key ran up $82,314.44 in 48 hours. Report cost per completed task beside quality, so savings never come from quietly worse answers.
Sources
- A postmortem of three recent issuesAnthropic Engineering, September 2025
- Model deprecationsAnthropic, Claude Developer Platform documentation
- Foundry Models lifecycle and support policyMicrosoft Learn, July 2026
- Silent Updates: Measuring and Closing the Post-Deployment Disclosure GapAbraham and Bucknall, arXiv 2608.11803, August 2026
- State of Agent EngineeringLangChain, survey of more than 1,300 professionals, December 2025
- Article 26: Obligations of Deployers of High-Risk AI SystemsEU Artificial Intelligence Act
- Article 73: Reporting of Serious IncidentsEU Artificial Intelligence Act
- Gartner: Half of Gen AI Projects Could Exceed Budget by 2028Campus Technology, June 22, 2026
- The 2026 AI Index Report: Responsible AIStanford HAI, 2026
- OpenAI pulls plug on overly supportive ChatGPT smarmbotThe Register, April 30, 2025
- API, ChatGPT and Sora facing issuesOpenAI status incident report, December 11, 2024
- Elevated error ratesOpenAI status incident, June 10, 2025
- Air Canada found liable for chatbot's bad advice on plane ticketsCBC News, February 2024
- NYC's AI Chatbot Tells Businesses to Break the LawThe Markup, March 29, 2024
- DPD chatbot goes off the rails at suggestion of customerThe Register, January 23, 2024
- Incident 622: Chevrolet Dealer Chatbot Agrees to Sell Tahoe for $1AI Incident Database
- How I Dropped Our Production Database and Now Pay 10% More for AWSAlexey Grigorev, DataTalks.Club, 2026
- Dev stunned by $82K Gemini API key bill after theftThe Register, March 3, 2026
- Klarna changes its AI tune and again recruits humans for customer serviceCX Dive, May 2025
- DeprecationsOpenAI API documentation
- Model lifecycleAmazon Bedrock User Guide
- Model lifecycle (models launched before September 7, 2026)Amazon Bedrock User Guide
- How Is ChatGPT's Behavior Changing over Time?Chen, Zaharia and Zou, Stanford and UC Berkeley, arXiv 2307.09009
- Amazon Bedrock Service Level AgreementAmazon Web Services
- What if your agent's hallucinations had a budget? How to start using SLOs for agent behaviorGrafana Labs, 2026
- Implementing SLOsGoogle, The Site Reliability Workbook
- GENOPS02-BP02 Monitor foundation model metricsAWS Well-Architected Generative AI Lens
- Observability in generative AIMicrosoft Learn, Microsoft Foundry, July 2026
- State of AI EngineeringDatadog, July 2026
- Semantic Conventions v1.42.0 release notesOpenTelemetry, June 2026
- Building effective agentsAnthropic Engineering, December 19, 2024
- Article 72: Post-Market Monitoring by Providers and Post-Market Monitoring Plan for High-Risk AI SystemsEU Artificial Intelligence Act
- Digital Omnibus on AIEU Artificial Intelligence Act explorer, 2026
- AI RMF Playbook: ManageNational Institute of Standards and Technology
- Commission Delegated Regulation (EU) 2025/301 on the content and time limits for reporting major ICT-related incidentsEUR-Lex, 2025
- Supervisory Letter SR 26-2 on Revised Guidance on Model Risk ManagementBoard of Governors of the Federal Reserve System, April 17, 2026
- Visual memo: Key changes under the federal banking agencies' revised model risk management guidanceDavis Polk, 2026
- Colorado Enacts New Law Regulating Automated Decision-Making TechnologyLathrop GPM, June 1, 2026
- Hidden Technical Debt in Machine Learning SystemsSculley et al., Google, NeurIPS 2015
- The State of AI in the Enterprise, 2026Deloitte, survey of 3,235 leaders
- Global Talent Shortage Reaches Turning Point as AI Skills Claim Top SpotManpowerGroup, February 26, 2026
- Gartner: Worldwide AI Spending to Reach $2.67 Trillion in 2026THE Journal, September 21, 2026
- GenAI Incident Response Guide 1.0OWASP GenAI Security Project, July 28, 2025
- Coalition for Secure AI Releases Two Actionable Frameworks for AI Model Signing and Incident ResponseOASIS Open, November 18, 2025
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 4, 2026. No client data appears in our insights.