AI agents Guide
AI SRE Agents in 2026: What They Find, Where They Fail and How to Deploy Them Safely
Every major cloud and observability vendor now sells an AI agent that investigates production incidents. Independent benchmarks show these agents narrow the search well and name the full root cause far less often than their marketing implies. This guide sets out what they find, where they fail, what they cost and how to give one access to production without adding a new way to break it.
For CTOs, heads of platform and operations, and SRE leaders deciding whether to buy, trial or limit an AI incident-investigation agent.
The short answer
AI SRE agents are useful for investigation: they read logs, metrics, traces and recent changes in minutes and hand engineers a ranked shortlist. On the hardest independent benchmarks they name the complete root cause in roughly a quarter to a half of incidents, so deploy them read-only, measure them on your own past incidents, and let any fix go through the same reviewed change path a human uses.
Key takeaways
- On realistic benchmarks, agents recover the exact root cause far less often than they find a suspect: 20.7% on average across 11 frontier models, while 76.0% name at least one correct service.1
- The best published score on IBM's ITBench-AA Kubernetes tasks is 56.2%, and on new, high-fidelity faults end-to-end success falls to between 15.4% and 28.2%.23
- Operators that published production results, including Alibaba and Meta, report faster investigations and useful shortlists. None reports an agent fixing incidents on its own.45
- Most products ship read-only by default, and where prices are public an investigation costs a few dollars. Every customer result we found is published by the vendor.678
- Outages come mostly from changes and automation. In the postmortems we reviewed, alerts fired within minutes and finding the cause took far longer.91011
- Reporting clocks start before the root cause is known: 6 hours under CERT-In in India, 4 hours from classification under the EU's DORA and 24 hours under NIS2.121314
An AI SRE agent is software that responds to a production alert by investigating it the way an on-call engineer would. It queries logs, metrics and traces, reads recent deployments and configuration changes, sometimes reads source code, and reports a likely root cause with its evidence. By October 2026, AWS, Microsoft, Google, Datadog, Grafana, Splunk, PagerDuty and a group of funded startups all sell one, and Gartner predicts that 85% of enterprises will use such tooling by 2029.15 The useful questions for a buyer are how often these agents are right, what they are allowed to touch, and how to measure one on your own incidents before trusting it.
- 20.7%of incidents in which frontier models recovered the exact set of root causes, averaged across 11 models1
- 56.2%top score on ITBench-AA's 59 Kubernetes incident tasks, as of October 20262
- 55.8%or more of coding-agent runs broke at least one action boundary when DevOps instructions were ambiguous16
What the benchmarks measured
Benchmarks with one injected fault and clean inputs produce flattering numbers. Benchmarks that require every root cause, start from a vague incident report or inject new kinds of failure produce much lower ones, and they are closer to what an on-call engineer faces. IBM's original ITBench found agents resolved 13.8% of SRE scenarios in early 2025.17 Its successor, ITBench-AA, launched in May 2026 with every frontier model below 50%;18 by October the top score was 56.2%, measured as average precision at full recall across 59 Kubernetes tasks.2
The failure mode matters more than the headline. In OpenRCA 2.0, agents named at least one correct root-cause service in 76.0% of cases but grounded it in a verified causal path to the symptom in only 61.5%, which the authors call an "ungrounded diagnosis."1 On ORCA-bench, the weakest model invented a root cause matching none of the valid ones in 40% of incident reports, against 7.2% for the most careful, and the authors note that real systems are "orders of magnitude larger, more dynamic, and more idiosyncratic" than any test bed.19 On Cloud-OpsBench, one widely used model reached a conclusion without calling a single diagnostic tool in 32% of cases.20 Model choice is a safety decision as well as an accuracy one.
Repair is harder still. SRE-Marathon runs about sixty overlapping faults through each 90-minute test, and the best of ten methods scored 41.3 out of 100. Agents "often correlate and localize faults, but almost never complete repairs while the faults remain active," and a fixed, non-AI runbook scored 33.5, ahead of most AI agents.21 Design moves results as much as the model does. One study of 3,500 diagnostic runs found that changing the underlying model moved accuracy more than five times as much as changing the agent framework, and that a verification step raised accuracy from 43.5% to 52.5%.22 Another audit found that the leading method on a pooled leaderboard picked the worse option on some subsystems, with a gap of up to 24.8 points,23 so a leaderboard winner is not necessarily the best agent for your stack.
What operators report from production
The companies that published results from their own incidents describe a narrower and more useful tool than the marketing does. Alibaba Cloud ran its BiAn system for 10 months on network failures and reduced time to find the faulty device by 20.5%, and by 55.2% for high-risk incidents.4 Meta's 2024 system narrows hundreds of candidate code changes to a list of five and achieved 42% accuracy for investigations in its web monorepo, and Meta warned it "can potentially suggest wrong root causes and mislead engineers."5 Meta's DrP platform, used by more than 300 teams, reduced time to resolve by 20 to 80%, largely through codified analyses written by engineers.24
Microsoft's figures are the most often misquoted. RCACopilot's "accuracy up to 0.766" measures how well it predicts an incident's root-cause category, not the specific cause.25 In Microsoft's study of 44,340 incidents, the owners of those incidents gave the best model's suggested root causes mean correctness scores of 2.88 and 2.56 out of 5.26 Across all the operator evidence, a human validates a ranked shortlist. No operator has published results from an agent resolving incidents without a person in the loop.
The products, what they may change and what they cost
| Product | Status, October 2026 | What it may change | Published price |
|---|---|---|---|
| AWS DevOps Agent | Generally available in six regions8 | No resource changes; it can open tickets and support cases27 | $0.0083 per agent-second6 |
| Azure SRE Agent | Available | Review mode proposes actions for approval; Autonomous mode acts within granted permissions28 | 4 Azure agent units per agent-hour always on, plus token use29 |
| Google Gemini Cloud Assist investigations | Since April 10, 2026, only for Premium Support customers or on request30 | Recommends; does not execute | Not published |
| Datadog Bits AI SRE | Generally available since December 2, 202531 | Investigations; actions through Datadog's workflow automation | About 6.5 AI credits per investigation; credits from $500 per 500 a month, or $1.30 on demand7 |
| Grafana Cloud Investigations | Available | Analysis and recommended next steps | From $2 per million tokens, "no per-seat or per-investigation fee"32 |
| Splunk AI SRE | Available; remediation in alpha33 | Remediation proposed as a pull request | Not published |
| PagerDuty SRE Agent | Fully autonomous responder in early access in the second half of 202634 | Diagnostics; autonomous response in early access | Part of platform tiers |
| incident.io AI SRE | Available | "The only change Investigations can make to your systems is a pull request you review and merge yourself"35 | Add-on |
| Cleric | Available | Investigations and change verification | Team plan $700 a month for 1,400 credits; 10 credits per investigation36 |
Where prices are public, an investigation costs a few dollars: about $6.50 to $8.45 at Datadog's credit rates,7 about $5 at Cleric's team rate,36 and $29.88 per agent-hour at AWS, so a ten-minute investigation costs about $5 before any query charges.6 Azure bills a fixed 4 units per agent-hour for as long as the agent exists, and when an agent reaches its monthly limit "it becomes unavailable for chat and actions until the next month,"29 a term worth negotiating before a month with many incidents. The larger costs are the observability platform the agent depends on and the engineering time spent checking answers that turn out to be wrong.
Data handling differs more than capability. AWS's agent "does not filter PII" from what it gathers, and it is "not impacted by customer policies in Service Control Policies (SCPs) or Control Tower that restrict customer content to specific regions," with some regions using global cross-region inference.27 Google says investigation data "can be stored in any Google Cloud data center" and advises against investigating data subject to residency rules.30 incident.io says it redacts sensitive data before it reaches model providers and has zero data retention agreements with them.35 For a company in India, where CERT-In requires 180 days of logs held within the country,12 where the agent's inference runs is a question for the contract.
The customer results in circulation are all published by vendors. AWS reports that Western Governors University cut resolution time on one disruption from an estimated two hours to 28 minutes.8 Resolve AI reports up to 87% faster time to root cause at DoorDash,37 and Traversal reports a 38% reduction in time to resolve at DigitalOcean.38 Each uses a different measure, none is published by the customer, and no independent comparison of commercial agents exists. The money is real all the same: Resolve AI raised at a $1.5 billion valuation in April 2026.39
What actually breaks production
| Outage | What triggered it | How long it took |
|---|---|---|
| Google Cloud, June 12, 2025 | A policy change reached new code that had no error handling and "nor was it feature flag protected" | Root cause identified within 10 minutes; fix rolled out within 40 minutes40 |
| AWS us-east-1, October 19 and 20, 2025 | "A latent race condition" in DynamoDB's DNS automation left an empty DNS record | Cause identified after 50 minutes; impact lasted from 11:48 PM to 2:20 PM PDT9 |
| Azure Front Door, October 29, 2025 | Valid customer configuration changes exposed "a latent bug in the data plane" | About eight and a half hours41 |
| Cloudflare, November 18, 2025 | A database permissions change doubled the size of a configuration file | Engineers "initially wrongly suspected" a DDoS attack; core traffic recovered after about three hours10 |
| Cloudflare, February 20, 2026 | A cleanup task withdrew 25% of customers' own IP prefixes | 1,100 prefixes withdrawn from 17:56 to 18:46 UTC42 |
| Azure West US, July 23, 2026 | A defect in the "blast radius analysis system" widened a repair to every optical device leaving a datacenter | Just under five hours43 |
Two lessons follow for AI agents. The slow step is diagnosis, which is where agents help, but Cloudflare's engineers first chased an attack because errors came and went, and an agent reading the same signals would face the same trap.10 And automation is itself a leading cause of large outages. AWS's DNS automation "required manual operator intervention to correct,"9 Facebook's 2021 outage passed because "a bug in that audit tool prevented it from properly stopping the command,"44 and Google's SRE book recounts automation that read an empty list of machines as "everything."45 Giving an agent write access adds one more automated actor to that list.
Independent survey data supports the picture. Uptime Institute's 2026 analysis finds that network problems cause the largest share of IT service outages and are "driven primarily by configuration and change management failures." In its 2025 survey, 57% of respondents put the cost of their latest major outage above $100,000, and the analysis warns that "AI could itself introduce errors."11 Google's SRE book reported in 2016 that "roughly 70% of outages are due to changes in a live system."46 The downtime prices that circulate come from vendor surveys: New Relic's median of $2 million an hour for high-impact outages47 and PagerDuty's $4,537 a minute.48 Engineers are less convinced of the payoff than their managers: in Catchpoint's 2026 SRE Report, 60% of directors said AI had reduced toil, against 38% of individual contributors.49
When AI agents act on their own
We found no primary-sourced case of an AI SRE agent causing a production outage. The cases on record involve coding agents with access broader than their task. In July 2025 a Replit agent deleted a production database during a code freeze and told its user that a rollback would not work, although the data was later recovered; Replit's chief executive called it "unacceptable and should never be possible."50 Amazon disputed a report that its Kiro agent caused an AWS outage, saying the event was "the result of user error, specifically misconfigured access controls" and affected a single service in one region.51 Research shows the general risk: when DevOps instructions were ambiguous, coding agents violated at least one action boundary in 55.8 to 67.8% of runs, and warnings about blast radius "barely reduce action propensity."16 Limits have to be enforced outside the model.
Reporting clocks start before the root cause is known
| Rule | First deadline | What starts the clock |
|---|---|---|
| India, CERT-In directions (2022) | 6 hours; logs kept 180 days in India | Noticing an incident, or being told of it12 |
| EU DORA, Delegated Regulation 2025/301 | 4 hours, and no later than 24 hours after becoming aware; intermediate report within 72 hours; final within a month | Classifying the incident as major13 |
| EU NIS2 Directive | 24-hour early warning; 72-hour notification; final report within a month | Becoming aware of a significant incident14 |
| UK Cyber Security and Resilience Bill | Proposed: 24-hour initial notice, 72-hour full report | Not yet law in October 202652 |
The tightest deadlines expire long before many investigations reach a confident root cause, which changes what an agent is for. Its regulatory value is assembling the timeline, the affected services and the evidence fast enough to classify an incident and start each clock correctly. An agent that records when the first signal fired, when a person became aware and when the incident was classified produces exactly the evidence regulators ask about.
How to give an agent access safely
No binding standard covers AI agents in operations yet. NIST launched an AI Agent Standards Initiative in February 2026.53 The practical guidance is OWASP's: its "Excessive Agency" risk names "excessive functionality; excessive permissions; excessive autonomy" as the causes, and its mitigations include least privilege, human approval for high-impact actions and authorization enforced in the systems the agent calls.54 OWASP's Top 10 for Agentic Applications extends this to tool misuse and privilege abuse.55 Engineers are more willing than the evidence supports: in Grafana's 2026 survey of 1,363 practitioners, autonomous actions had 77% support, and alert fatigue was the biggest single obstacle to faster incident response, cited by 30%.56
| Level | What the agent may do | When to grant it |
|---|---|---|
| Read | Query logs, metrics, traces, change history and code under its own short-lived identity | From day one, after redaction and region checks |
| Propose | Write a fix as a pull request, a runbook step or a change request a person approves | After it beats your baseline on past incidents |
| Act on a narrow path | Run named, reversible actions such as a rollback or a restart, with an action log | Only where the action is already automated and tested |
| Act broadly | Change production on its own judgment | Not supported by current evidence |
Sending logs to a model provider is also a data-protection decision, because logs carry IP addresses, user IDs and email addresses. Anthropic's own API, for example, offers US-only or global inference, with other regions available through cloud providers' endpoints.57 Check where inference runs, what is retained and whether personal data is redacted before it leaves your collectors.
Read-only investigator
How we advise, by default
- Shortlist and evidence in minutes
- No new path to change production
- Timeline for reporting clocks
- Engineers decide and act
Proposes fixes for approval
Once it beats your baseline
- Fixes arrive as pull requests or change requests
- Same review as a human change
- Faster recovery on known failures
- Needs a reviewer on call
Autonomous remediation
Vendor roadmaps, early access
- Acts within granted permissions
- Repair rare in independent tests
- Adds an automated actor to production
- Hard to audit after the fact
How to trial an AI SRE agent
- Pick ten past incidents Choose recent postmortems with known root causes, including at least two messy ones with several causes or a misleading first signal.
- Set the baseline Record how long your team took to find each cause and what the first wrong lead was.
- Check the data path Confirm where inference runs, what is stored, how personal data is redacted and whether the vendor's terms meet your reporting and residency rules.
- Grant read access only Give the agent its own identity with read scopes on telemetry, change history and code, and no write permissions.
- Score the evidence Replay each incident and score the causal chain the agent shows, not only the service it names; count invented causes separately.
- Run it beside on-call For 30 to 60 days, let it investigate live alerts alongside engineers and measure time to a correct cause.
- Add proposals, not actions If it beats the baseline, let it propose fixes as pull requests or change requests through your normal review.
Questions before buying an AI SRE agent
- How does it score on our own past incidents, judged on the evidence chain?
- What can it change in production today, and what permissions does it ask for?
- Where does inference run, what is retained and how is personal data redacted?
- What does an investigation cost, and what happens when a usage cap is reached?
- Can it produce the timeline our reporting deadlines require?
- Who reviews its proposals, and how are its actions logged and reversed?
The evidence moves the buying question from "can the agent find the root cause?" to "how much faster does it make our people, and what is it allowed to touch?" Agents shorten the correlation work that dominates real outages and stay unreliable on the incidents that hurt most: new faults, interacting failures, vague reports and busy systems. That makes an AI SRE agent a fast investigator for your engineers. Measure it on your own incidents, keep it read-only until the numbers say otherwise, and let any change it proposes travel the same reviewed path as a human's.
This is how we approach incident response in our managed services work: replay past incidents to set a baseline, give investigation agents read-only access under their own identity, keep fixes in the reviewed change path, and build the incident timeline that reporting rules ask for.
Questions leaders ask
What is an AI SRE agent?
An AI SRE agent is software that investigates production incidents the way an on-call site reliability engineer would. It queries logs, metrics and traces, reads recent changes and sometimes source code, and reports a likely root cause with evidence. Some products can also propose or carry out fixes, usually behind human approval.
How accurate are AI SRE agents at finding root causes?
It depends on how hard the incident is. On the hardest 2026 benchmarks, the best agents recover the complete root cause in roughly a quarter to a half of incidents, and the average across 11 frontier models on OpenRCA 2.0 was 20.7%. Agents name at least one correct service far more often, in 76.0% of cases on that benchmark.
Can AI agents fix production incidents on their own?
Rarely, in the evidence so far. On SRE-Marathon, agents almost never completed repairs while faults were still active, and a fixed runbook beat most of them. Operators that publish results use agents to shorten investigations, with engineers making the fix. Most products ship read-only or route fixes through pull requests and approval.
How much does an AI SRE agent cost?
Where vendors publish prices, an investigation costs a few dollars: AWS charges $0.0083 per agent-second, Datadog's investigations use about 6.5 credits, and Cleric charges 10 credits per investigation. The larger costs are the observability platform the agent relies on and the time engineers spend checking its answers.
Is it safe to give an AI agent access to production?
Read access under the agent's own identity, with personal data redacted and inference in an approved region, is a reasonable start. Write access should come through the same reviewed change path a person uses. Research found coding agents broke action boundaries in more than half of runs when instructions were ambiguous.
What is the difference between AIOps and an AI SRE agent?
AIOps usually refers to machine learning that groups alerts, detects anomalies and reduces noise. An AI SRE agent goes further: it runs an investigation across tools, reasons over the results and writes up a likely root cause, much as an engineer would.
How do incident-reporting rules affect AI incident response?
The first deadlines are short: 6 hours under India's CERT-In directions, 4 hours from classification under the EU's DORA, and 24 hours under NIS2. They usually expire before a root cause is confirmed, so an agent's main regulatory value is assembling the timeline and impact evidence quickly.
Sources
- OpenRCA 2.0arXiv 2606.27154, June 2026
- ITBench-AA leaderboardArtificial Analysis and IBM Research
- SREGym: a benchmark for AI SRE agents on live clustersarXiv 2605.07161, May 2026
- BiAn: LLM-based network failure localization at Alibaba CloudAlibaba, SIGCOMM 2025
- Leveraging AI for efficient incident responseEngineering at Meta, June 2024
- AWS DevOps Agent pricingAWS
- Bits AI pricingDatadog
- Announcing general availability of AWS DevOps AgentAWS, March 2026
- Summary of the Amazon DynamoDB service disruption in the Northern Virginia (US-EAST-1) RegionAWS, October 2025
- Cloudflare outage on November 18, 2025Cloudflare
- Annual outage analysis 2026Uptime Institute, 2026
- Directions under section 70B of the IT Act, April 28, 2022CERT-In
- Commission Delegated Regulation (EU) 2025/301 on reporting major ICT-related incidentsEUR-Lex
- Directive (EU) 2022/2555 (NIS2)EUR-Lex
- Gartner Market Guide for AI SRE ToolingGartner, January 2026, via Cast AI
- Action boundaries of coding agents under ambiguous DevOps instructionsarXiv 2607.02294, July 2026
- ITBench: evaluating AI agents across diverse real-world IT automation tasksJha et al., IBM, arXiv 2502.05352, 2025
- ITBench-AA: a benchmark for AI agents in IT operationsIBM Research, Hugging Face blog, May 2026
- ORCA-bench: root cause analysis on realistic incident reportsarXiv 2607.28545, 2026
- Cloud-OpsBencharXiv 2603.00468, 2026
- SRE-Marathon: agents under continuous, overlapping faultsarXiv 2609.33023, September 2026
- A study of 3,500 RCA agent trajectoriesarXiv 2608.21310, 2026
- An audit of root cause analysis leaderboardsarXiv 2606.29159, 2026
- DrP: Meta's root cause analysis platform at scaleEngineering at Meta, December 2025
- Automatic root cause analysis via large language models for cloud incidents (RCACopilot)Chen et al., Microsoft, EuroSys 2024
- Recommending root-cause and mitigation steps for cloud incidents using large language modelsAhmed et al., Microsoft, ICSE 2023
- AWS DevOps Agent securityAWS documentation
- Run modes in Azure SRE AgentMicrosoft Learn
- Billing for Azure SRE AgentMicrosoft Learn
- Gemini Cloud Assist investigationsGoogle Cloud documentation
- Datadog launches Bits AI SRE agentDatadog, December 2025
- Grafana Cloud InvestigationsGrafana Labs
- AI SRE in Splunk Observability Cloud can remediate incidentsSplunk, September 2026
- PagerDuty Operations Cloud spring 2026 releasePagerDuty, 2026
- AI SREincident.io
- PricingCleric
- DoorDash customer storyResolve AI
- DigitalOcean customer storyTraversal
- Resolve AI announces $40 million Series A extension at $1.5 billion valuationGunderson Dettmer, April 2026
- Multiple Google Cloud products experiencing service issues, June 12, 2025Google Cloud incident report
- Post incident review: Azure Front Door, October 29, 2025Microsoft Azure status history
- Cloudflare outage on February 20, 2026Cloudflare
- Post incident review: Azure West US, July 23, 2026Microsoft Azure status history
- More details about the October 4 outageEngineering at Meta, October 2021
- Automation at GoogleSite Reliability Engineering, Google
- IntroductionSite Reliability Engineering, Google, 2016
- Observability forecast 2025New Relic, September 2025
- Study on the cost of incidentsPagerDuty, 2024
- SRE Report 2026Catchpoint, 2026
- AI coding tool Replit wiped a database and called it a catastrophic failureFortune, July 2025
- AWS service outage and the Kiro AI toolAbout Amazon
- Cyber Security and Resilience Bill factsheet: incident reportingGOV.UK, 2026
- Announcing the AI Agent Standards InitiativeNIST, February 2026
- LLM06:2025 Excessive AgencyOWASP GenAI Security Project
- OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project, December 2025
- Grafana Labs' 4th annual Observability SurveyGrafana Labs, March 2026
- Data residencyClaude Developer Platform documentation
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 6, 2026. No client data appears in our insights.