AI agents Guide

Computer-Use Agents vs RPA in 2026: What AI That Operates Screens Can Replace

AI agents can now read a screen, move the mouse and type into almost any application, which makes them look like a replacement for every RPA bot a company runs. The best of them score above 85% on short desktop benchmarks. On long workflows, repeated runs and real business software, the numbers fall away fast. This guide sets out what the evidence says agents can take over, what bots still do better, and how to route the work between them.

For COOs, CIOs, heads of shared services and automation leads deciding whether to replace, extend or keep their RPA estate as computer-use AI agents arrive.

Published
Reviewed
Reading time
19 min

The short answer

Computer-use agents can take over screen work that RPA never paid for: low-volume tasks, screens that change often and systems without an API. They cannot yet replace bots on stable, high-volume processes, where a bot is faster, cheaper per run and gives the same result every time. Route each step to an API, a bot, an agent or a person, and check every agent write against the system of record.

Key takeaways

  • Short benchmarks are nearly saturated, with self-reported OSWorld-Verified scores up to 86.1%, while the best published strict score on the long workflows of OSWorld 2.0 is 48.7%.12
  • On real ERP software one agent saved records in up to 85% of runs but wrote the correct value in as few as 3%, an error a screenshot does not show.3
  • Repetition is the weak point: a strong agent solved about 78% of OSWorld tasks at least once in ten tries and all ten times on only about 36%.4
  • In the one controlled test against UiPath, the bot ran 10 out of 10 times on every task and the agent was slower and less reliable, but took minutes to build where the bot took hours.5
  • Microsoft charges 5 Copilot Credits for a computer-use step and 13 credits per 100 deterministic flow actions, about 38 times more per step.6
  • Prompt injection is down to fractions of a percent in vendor tests, and OpenAI and the UK NCSC both say it may never be fully solved.278

A computer-use agent is an AI model that operates software the way a person does. It takes a screenshot, decides what to click or type, acts, and looks again, until the task is done. Robotic process automation (RPA) also works through the screen, but a bot follows steps a developer recorded in advance and breaks when a field moves. The agent decides its path at run time, so it copes with changing layouts and messy input, and it can also choose a wrong path. Since 2025 every major model vendor and every RPA vendor has shipped a version of computer use, and the question for anyone running a bot estate is which work moves, which stays, and how to prove the agent did it right.

  • 36%of OSWorld tasks a strong agent completed on all ten of ten runs, against about 78% completed at least once4
  • 3%correct values written by some agents on ERP tasks where they saved the record in up to 85% of runs3
  • 38xthe cost of a computer-use step against a deterministic flow action in Microsoft Copilot Studio6

What the benchmarks measure and what they leave out

The best-known test is OSWorld, a set of desktop tasks in a virtual machine. On its cleaned-up version, OSWorld-Verified, the top 2026 entries run from 83.4% to 86.1%, and every one is reported by the vendor that built the model. The leaderboard that collects them warns that rows "can vary by evaluator, harness, attempt budget, tool access, task filtering, or verification level".1 The human score usually set beside them, 72.36%, was measured on the original 2024 task set.9 Stanford's AI Index 2026 put agents at 66.3% on OSWorld, "within 6 percentage points of human performance", in data that is already six months old.10

Longer work tells a different story. OSWorld 2.0, released in June 2026, has 108 workflows that take a person a median of about 1.6 hours and an agent an average of 318 tool calls, against about 30 in the original. At launch the best agent completed 20.6% of them.11 Anthropic's system card for Claude Opus 5.5 reports 48.7% completed on a strict measure and 81.8% on partial credit.2 Tests built on real business applications show the same drop as tasks get longer.

Exhibit 1Best published agent results, short tasks against long workflows (%)
  • OSWorld-Verified desktop tasks, self-reported86.1%
  • ERPBench single ERP record94%
  • UI-CUBE simple enterprise UI actionsup to 85%
  • OSWorld 2.0 workflows, strict48.7%
  • UI-CUBE complex enterprise workflowsup to 19%
  • SaaS-Bench tasks across business apps3.8%
Each benchmark uses its own tasks and scoring, so compare the shape, not the rows against each other. SaaS-Bench counts a task only when every checkpoint passes. Sources: [1], [3], [12], [2], [13]
BenchmarkWhat it testsBest resultHuman reference
ERPBench (Accenture, Sept 2026)30 tasks in the open-source ERPNext, scored on database valuesClaude Sonnet 4.6 completed 94% of single-record tasks and 100% of the two harder tiers; the strongest open models reached 34% and 32% on single records387% to 100% across three annotators
LegacyWorld (Aug 2026)28 legacy Windows business and healthcare workflowsClaude Opus 4.6 78.6% valid success; GPT-5.4 3.6%14Not published
UI-CUBE (UiPath, Nov 2025)226 enterprise UI tasks67% to 85% on simple actions, 9% to 19% on complex workflows1297.9% simple; 61.2% complex for people new to the apps
SaaS-Bench (May 2026)106 tasks across 23 open-source business applications3.8% fully resolved by Claude Opus 4.713Not published
These are the closest public tests to back-office work. None of them runs on SAP GUI, Workday, Oracle E-Business Suite or mainframe terminals, where most RPA estates earn their keep.

Two lessons follow for buyers. A high general score does not predict results on your software: in ERPBench, Holo3-35B-A3B and Qwen3-VL-32B completed 34% and 32% of the simplest ERP tasks while Claude completed 94%.3 And the drop comes with length. UiPath's own researchers describe "a sharp capability cliff rather than gradual performance degradation" between simple and complex tasks.12 We found no independent public test of computer-use agents on SAP, Workday or mainframe screens, so any claim about those systems needs your own trial.

Why one successful run is the wrong measure

A bot that posts invoices runs the same process thousands of times a month, so what matters is whether it succeeds every time. A 2026 study ran a strong agent, Agent S3 with GPT-5, ten times on each OSWorld task. It completed about 78% of tasks at least once, and all ten times for only about 36%.4 SaaS-Bench saw one model score anywhere from 0 to 0.679 on the same task across runs, and set out the arithmetic of long work: at 95% accuracy per checkpoint across 12 checkpoints, the whole task succeeds only about 54% of the time.13

Most failures are failures of judgment. An audit of failed runs across five benchmarks traced 35.2% of genuine failures to planning and 39.3% to verification and feedback, with the largest single cause, at 29.5%, an agent repeating an action that did nothing. Clicking the wrong thing accounted for 13.9%.15 The failure that matters most to a finance team is the silent one. In ERPBench some agents saved the record in up to 85% of runs and wrote the correct value in as few as 3%.3 In LegacyWorld, Claude Opus 4.6 left unwanted changes behind in 10.7% of runs and Kimi K2.5 in 35.7%.14 SaaS-Bench documents agents claiming success after their own verification failed.13

Benchmarks carry their own error. The same audit found 15.3% of failure verdicts were wrong, 17.5% on OSWorld and 21.7% on WebArena.15 Agents are also slow. The best take 2.7 to 4.3 times more steps than necessary, and a late step can take three times longer than an early one.16 In ERPBench, Claude Sonnet 4.6 spent 89 seconds and 231,000 input tokens on a single record, and 323 seconds and about 2 million input tokens on a chained workflow.3 At today's $2 per million input tokens for Claude Sonnet 5.5, that is roughly $0.46 and $4 a run before caching, by our arithmetic.17

The one controlled test against an RPA bot

Researchers at the Technical University of Liberec built the same three processes in UiPath and with Anthropic's computer-use agent on Claude Sonnet 4, then ran each ten times. The tasks came from rpachallenge.com, a set of exercises the RPA community uses to test bots.5

TaskUiPath botComputer-use agentTime to build
Copy spreadsheet rows into a web form whose layout changes each round139.8 seconds, 10 of 10 runsDid not complete its one runAbout 40 minutes for the bot; the agent never reached a working version
Watch a stock price and alert below a threshold53.9 seconds, 10 of 10109.8 seconds, 9 of 10About 38 minutes against about 10
Read invoices and enter their data20 seconds, 10 of 10202.8 seconds, 6 of 10About 240 minutes against about 15
Průcha, Matoušková and Strnad, arXiv 2509.04198, a 2025 preprint with one developer and small samples. One agent run of the invoice task cost about $0.28.

The authors call the agent "not yet production-ready" and found it far quicker to build.5 The bot was faster on every task, and the differences in reliability were too small a sample to be statistically significant. The study used a 2025 model, and later benchmarks point the same way: agents trade run time and repeatability for build time and tolerance of change. That trade decides where each one belongs.

What computer use costs

Model vendors now bill computer use as ordinary tokens, and the buyer runs the browser or virtual machine. Microsoft and Amazon sell it as a metered unit. RPA vendors sell bots by the month and agent calls by platform credit.

ProductStatus, October 2026Published priceWhere it runs
OpenAI computer use in the Responses APICurrent, replacing computer-use-preview18GPT-6 Astra $10 input and $50 output per million tokens; GPT-6.1 Sol $2 and $1019Your browser or virtual machine
Anthropic computer use toolClaude API and Google Cloud; beta on Amazon Bedrock, Microsoft Foundry and Claude Platform on AWS20Claude Sonnet 5.5 $2 and $10; Opus 5.5 $4 and $20; the tool adds about 4,500 input tokens per request17Your virtual machine or container
Google Gemini API computer useGemini 3.5 to 3.8 Flash models21Gemini 3.8 Flash $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.5022Your client environment
Microsoft Copilot Studio computer useOpenAI's agent and Claude Sonnet 4.5 generally available; Claude 4.6 models experimental235 Copilot Credits a step, 15 on premium models; $200 for 25,000 credits, about $0.04 or $0.12 a step; not included in Microsoft 365 Copilot user licenses624A Windows machine
Amazon Nova ActAvailable on AWS$4.75 per agent hour of elapsed working time25AWS
Microsoft Power Automate (RPA)Current$15 per user a month; $150 per unattended bot a month; $215 per hosted bot a month26Your Windows machines or Microsoft-hosted
UiPathCurrentBasic from $25 a month, higher tiers by quote; each agent model call costs 0.16 to 0.4 Platform Units, counted per 64,000 input tokens2728UiPath cloud or your own servers
List prices checked on October 6, 2026. Automation Anywhere, SS&C Blue Prism and ServiceNow publish no list prices.

Microsoft is the one vendor that prices a bot, a deterministic flow and a computer-use step side by side, which makes it a clean comparison. A flow action costs 13 credits per 100 actions and a computer-use step costs 5 credits, about 38 times more.6 One hosted Power Automate bot at $215 a month buys the same as about 26,900 credits, or roughly 5,400 standard computer-use steps.2624 At 30 steps a task, that is about 180 agent tasks a month. The arithmetic leaves out build and maintenance labor, which is most of what RPA costs, but it shows where the line falls. On a stable screen running several hundred times a month, the bot is cheaper to run. On a screen that runs a few dozen times a month or changes every quarter, the agent's shorter build and lower upkeep can win.

How RPA has actually performed

The most quoted statistic about RPA, that 30% to 50% of projects fail, comes from a 2016 EY paper, where it is an observation about first projects: "we have seen as many as 30 to 50% of initial RPA projects fail".29 There is no survey behind it. The better evidence points at upkeep, scale and measurement. HFS Research estimated that licenses are "just 25% to 30% of total costs for implementing RPA".30 Deloitte's 2020 survey found 37% of organizations piloting with 1 to 10 automations and 13% scaling beyond 50.31 Its next survey found the average payback period for those still piloting had grown from 16 months in 2020 to 22 months.32

US government auditors give the clearest record. The inspector general for the General Services Administration found that the agency's claim that its RPA program reclaimed more than 240,000 work hours a year was "inaccurate and unreliable", and that GSA was not tracking what its bots cost.33 A 2024 follow-up counted 119 active and 24 decommissioned bots, found no process for removing retired bots' access, and found that program management "simply removed or modified the requirements" it could not meet.34 HUD's inspector general concluded that after more than three years its program "had achieved minimal progress and results".35 Measurement, cost tracking and machine identities are the same three problems agents bring, with less predictable behavior on top.

The RPA business itself is growing slowly. UiPath reported $1.938 billion of annual recurring revenue at July 31, 2026, up 12%, with $37 million net new.36 Eighteen of its top 20 deals that quarter included AI, and its chief executive described the pitch plainly: AI "is probabilistic and can be expensive at scale", while many processes "need exactness, the same result every time".37 Automation Anywhere says AI accounts for nearly 70% of its new and upsell bookings and that agent executions grew five times in a year.38 Analysts expect slow adoption: Forrester predicted that fewer than 15% of firms will turn on the agentic features in their automation suites in 2026,39 and Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027.40

What deployments look like so far

We found no company that has publicly retired a set of RPA bots in favor of agents and published the before and after costs. The public cases follow two patterns. The first is computer use where no API exists. Graebel, a relocation company, uses Copilot Studio computer use because its globalCONNECT platform "doesn't expose APIs for this workflow". The agent reads service orders, works through the platform's screens and routes "low-confidence cases, exceptions, and approvals for human review". Microsoft describes it as a pilot covering "a small set of high-volume service order types" and publishes no figures.41

The second pattern is agents working alongside existing bots, which is how the RPA vendors sell. UiPath told investors it handles about 700,000 invoices a year for a Fortune Global 500 manufacturer, and reported 96% document accuracy in a proof of concept.37 Amazon says Nova Act reached 90% reliability on browser workflows built by early customers, and that one startup client automates hundreds of thousands of workflows a month with it.42 These are vendor figures without independent measurement. In outsourced back-office work the effect so far shows up in revenue mix: Genpact's core business services grew 1.9% in its second quarter of 2026 while its technology business grew 24.1%.43

Prompt injection is lower in vendor tests and still open

A computer-use agent reads untrusted content, web pages, emails, documents and screenshots, and acts on it with the access it was given. Text planted in that content can redirect it. Vendor measurements have fallen fast. In August 2025 Anthropic measured a 23.6% attack success rate for browser use without its mitigations and 11.2% with them.44 A year later, on a harder set of attacks written by professional red-teamers, attacks succeeded against Claude Opus 4.5 17.6% of the time and against Opus 5 3.8% of the time before extra safeguards.45 The Claude Opus 5.5 system card reports 0.07% in computer-use environments, two successful attempts out of 2,800. The same card reports 54.61% for an adaptive attacker in coding environments, falling to 11.13% with probes enabled, which shows how much the number depends on the setting.2

No one who builds these systems calls the problem solved. OpenAI has said its AI browser may always be vulnerable to prompt injection.7 The UK's National Cyber Security Centre said such attacks "may never be totally mitigated" the way SQL injection can be, and told developers to focus on secure design.8 Independent researchers keep finding working attacks. Brave showed in August 2025 that a hidden comment on Reddit could make Perplexity's Comet browser log into the user's account, read a one-time password from Gmail and send both to the attacker.46 In academic tests, visual prompt injections fooled computer-use agents up to 51% of the time and browser agents up to 100% on some platforms,47 and WASP found attacks partly succeeding in up to 86% of cases while agents completed the attacker's goal in 0 to 17%, which its authors call "security by incompetence".48 That margin shrinks as agents get better at finishing tasks.

RiskControlWhat an auditor can check
Instructions planted in pages, emails or screenshotsIsolated virtual machine, domain allowlist, no standing credentials in the sessionThe machine image, the allowlist and the credential vault's access log
A wrong value saved to a system of recordA check against the database or API after every write, with mismatches sent to a queueCheck results per run and the mismatch queue
Payments, purchases, deletions and messages sent outsideA person confirms before the actionApproval records with the approver's identity
Access wider than the taskA separate least-privilege identity for each agent, reviewed and removed on retirementThe identity inventory and review dates
Behavior nobody can reconstructScreenshots and actions recorded for every run and keptSession logs tied to transaction numbers
OWASP's Top 10 for Agentic Applications, published in December 2025, puts agent goal hijack first.49 Anthropic's documentation recommends a dedicated virtual machine with minimal privileges, a domain allowlist and human confirmation for consequential actions.20
Rule or caseStatus, October 2026What it means for computer-use agents
Amazon v. Perplexity, US Court of Appeals for the Ninth CircuitInjunction against Perplexity's Comet agent vacated on August 4, 2026; rehearing by the full court declined5051Under US anti-hacking law it is the user who "accesses" a site through an agent. Contract, unfair competition and other claims remain open, so a portal's terms bind the company directing the agent
Web Bot Auth for signed agentsUsed by Visa, Mastercard and American Express for agent commerce through Cloudflare52Sites are starting to admit agents that prove who they are. Unsigned screen automation will meet more blocking on third-party sites
EU AI Act Article 50Most transparency duties apply from August 2, 2026; Annex III high-risk duties moved to December 2, 202753An agent that writes to people, such as suppliers or customers, must not pass as a person. Agents in hiring or credit decisions fall under the later high-risk regime
UK GDPR Articles 22A to 22DIn force since February 5, 202654Solely automated decisions with significant effects are allowed more widely in the UK, with safeguards. The EU rules on such decisions remain stricter
US bank model risk guidance, SR 26-2Replaced SR 11-7 on April 17, 202655Generative and agentic AI "are not within the scope of this guidance", so banks must set their own controls for agents
Proposals and draft standards are left out. NIST's AI standards center opened a request for information on securing AI agent systems in January 2026 and has not issued a standard.56

Which mechanism for which step

The decision is made step by step, and one process usually ends up mixing an API call, a bot, an agent and a person. The order of preference that the evidence supports is an API first, a deterministic bot second, an agent third, and a person for anything that cannot be checked automatically.

SituationUseWhy
The system has a documented API or connectorAPI or integration platformDeterministic, cheapest per run and fully logged
Stable screen, hundreds of runs a month or more, structured input, no APIRPA botFaster and repeatable; at volume the bot's monthly fee is a fraction of the agent's steps
Low volume, many different screens, layouts that change, no APIComputer-use agent with a check after every writeMinutes to build and tolerant of layout changes
Emails, PDFs or free text feeding a structured systemA model reads and classifies; an API or bot writesKeeps the uncertain step away from the write
Exceptions from an existing botAgent on the exception queue, bot on the main pathJudgment where it adds value, exactness where the volume is
Long workflow across several applicationsSplit into stages with a state check between eachEnd-to-end success on SaaS-Bench was 3.8%
SAP, mainframe terminals, Workday, OracleKeep the API or bot; trial an agent beside itNo independent public results on these systems
Decisions with legal or similar effects on a personA person decides, the agent preparesGDPR and the UK safeguards require meaningful human involvement
Exhibit 2Three ways to automate a screen task

API or integration

When the system offers one

  • Same result every time
  • Cheapest per run
  • Survives screen redesigns
  • Needs the vendor to expose the operation

RPA bot

Stable screens at volume

  • Runs in seconds
  • Predictable cost at high volume
  • Breaks when fields move
  • Hours to build and to repair

Computer-use agent

The long tail and changing screens

  • Minutes to set up
  • Copes with layout changes and messy input
  • Slower and varies between runs
  • Needs a check after every write
Most automation estates will run all three, with people on the exceptions.

How to start in eight weeks

  1. List the screen work Inventory every bot and every manual screen task with its monthly volume, how often the screens change, and whether the system has an API.
  2. Pick the long tail Choose three to five low-volume tasks with no API and frequent screen changes, which bots never covered economically.
  3. Define correct For each task, write the database or API check that proves the work was done right, independent of what the agent reports.
  4. Build the sandbox Give the agent its own virtual machine, its own least-privilege identity, a domain allowlist and no stored passwords.
  5. Run it ten times Run every task repeatedly on real cases and count how often it succeeds on all runs, alongside the time, steps and cost per run.
  6. Price at real volume Compare the cost per task at your monthly volume, including review time for exceptions, with a bot or a person doing the same work.
  7. Go live with a queue Send low-confidence cases and failed checks to people, record every session, and widen the scope only when the numbers hold.

Questions before replacing a bot with an agent

  • Does this system have an API we could use instead of its screens?
  • How many times a month does this task run, and how often do its screens change?
  • How will we prove each write was correct without trusting the agent's own report?
  • How often does the agent succeed on every one of ten runs on our real cases?
  • What does a run cost at our volume, including people reviewing exceptions?
  • What can the agent's identity reach if planted text redirects it?
  • Do the terms of any third-party portal it uses allow automated access?

Clicking is close to solved in 2026. Doing a long task correctly every time, and proving it, is not, and that is the thing RPA's rigid design was built to guarantee. Agents earn their place where bots were always too expensive to build and maintain: the low-volume screens, the changing layouts and the messy input that stayed manual. The organizations that gain most will be the ones that write the checks, measure repeated runs and route each step to the cheapest mechanism that can be trusted with it.

This is how we approach screen automation in our AI agent development work: inventory the bot estate and the manual screen work, use APIs wherever they exist, give agents the long tail inside an isolated sandbox with their own identity, and put a database check and an exception queue behind every write.

Questions leaders ask

Will AI agents replace RPA?

Not on stable, high-volume processes in 2026. In the only controlled comparison, a UiPath bot was faster and succeeded on every run, while the computer-use agent was slower and less reliable. Agents extend automation to low-volume screens, changing layouts and unstructured input that bots never covered economically, and they increasingly work alongside existing bots.

What is a computer-use agent?

An AI model that operates software through its screen. It takes screenshots, decides what to click or type, and repeats until the task is done. OpenAI, Anthropic, Google, Microsoft and Amazon all offer one, and RPA vendors such as UiPath and Automation Anywhere have added agents to their platforms.

How reliable are computer-use agents in 2026?

Good on short tasks and weak on long ones. The best agents score above 80% on short desktop benchmarks and 94% on single ERP records, but only 48.7% on 1.6-hour workflows and 3.8% on tasks across several business applications. Repeated runs matter most: one strong agent solved all ten runs on only about 36% of OSWorld tasks.

How much does a computer-use agent cost per task?

It depends on the steps and the model. Microsoft charges about $0.04 a step on standard models in Copilot Studio, Amazon charges $4.75 per agent hour for Nova Act, and model vendors bill tokens. A chained ERP task that used about 2 million input tokens would cost roughly $4 at $2 per million before caching.

Is it safe to give a computer-use agent company credentials?

Only with the access limited to what the task needs. Run the agent in its own virtual machine with its own least-privilege identity, a domain allowlist and no stored passwords, and require a person to confirm payments, deletions and outside messages. Vendors have cut prompt injection rates sharply, and none of them says the problem is solved.

Should we use computer use when the system has an API?

No. An API call gives the same result every time, costs far less per run and survives screen redesigns. Computer use is for systems without an API, or where the API does not expose the operation you need.

Can computer-use agents automate SAP or mainframe screens?

There are no independent public results for SAP GUI, Workday, Oracle E-Business Suite or mainframe terminals as of October 2026. If an API or a working bot exists, keep it, and trial an agent beside it on your own cases, with a database check behind every write.

Sources

  1. OSWorld-Verified leaderboardSteel.dev, 2026
  2. Claude Opus 5.5 System CardAnthropic, 2026
  3. ERPBenchAccenture, arXiv 2609.17885, September 2026
  4. On the reliability of computer use agentsarXiv 2604.17849, April 2026
  5. RPA and LLM agents compared on enterprise tasksPrůcha, Matoušková and Strnad, arXiv 2509.04198, September 2025
  6. Billing rates and managementMicrosoft Learn
  7. OpenAI says AI browsers may always be vulnerable to prompt injection attacksTechCrunch, December 2025
  8. Mistaking AI vulnerability could lead to large-scale breachesUK NCSC, December 2025
  9. OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environmentsXLANG Lab
  10. AI Index Report 2026: technical performanceStanford HAI, 2026
  11. OSWorld 2.0arXiv 2606.29537, June 2026
  12. UI-CUBEUiPath, arXiv 2511.17131, November 2025
  13. SaaS-BencharXiv 2605.15777, May 2026
  14. LegacyWorldarXiv 2608.14131, August 2026
  15. How benchmarks mis-score computer-use agentsarXiv 2607.28367, July 2026
  16. OSWorld-Human: benchmarking the efficiency of computer-use agentsarXiv 2506.16042, June 2025
  17. Claude API pricingAnthropic
  18. Computer useOpenAI API docs
  19. API pricingOpenAI
  20. Computer use toolAnthropic docs
  21. Computer useGoogle Gemini API docs
  22. Gemini Developer API pricingGoogle
  23. Automate web and desktop apps with computer useMicrosoft Learn
  24. Copilot Studio pricingMicrosoft
  25. Amazon Nova pricingAWS
  26. Power Automate pricingMicrosoft
  27. PricingUiPath
  28. Agents licensingUiPath docs
  29. Get ready for robotsEY, 2016
  30. How to make sense of nonsensical RPA software pricingHFS Research, 2018
  31. Automation with intelligence: 2020 surveyDeloitte, 2020
  32. Intelligent automation 2022 survey resultsDeloitte, 2022
  33. GSA's robotic process automation program lacks evidence to support claimed savingsGSA Office of Inspector General, November 2023
  34. GSA's robotic process automation program: security auditGSA Office of Inspector General, August 2024
  35. HUD's robotic process automation program was not efficient or effectiveHUD Office of Inspector General, 2023
  36. UiPath reports second quarter fiscal 2027 financial resultsUiPath, September 2026
  37. UiPath Q2 2027 earnings call transcriptThe Motley Fool, September 2026
  38. Automation Anywhere reports continued double-digit growth in Q2 FY27Automation Anywhere, September 2026
  39. Predictions 2026: automation at the crossroadsForrester, 2025
  40. Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027Gartner, June 2025
  41. Graebel agent transformation storyMicrosoft
  42. AWS agentic AI announcementsAbout Amazon, December 2025
  43. Genpact reports second quarter 2026 resultsGenpact, August 2026
  44. Piloting Claude for ChromeAnthropic, August 2025
  45. Claude in Chrome is generally availableAnthropic, August 2026
  46. Agentic browser security: indirect prompt injection in Perplexity CometBrave, August 2025
  47. VPI-Bench: visual prompt injection attacks for computer-use agentsarXiv 2506.02456, 2025
  48. WASP: benchmarking web agent security against prompt injection attacksarXiv 2504.18575, 2025
  49. OWASP Top 10 for Agentic ApplicationsOWASP GenAI Security Project, December 2025
  50. Ninth Circuit holds human user, not developer of AI agent, responsible for websites accessedTroutman Pepper Locke, August 2026
  51. 9th Circuit won't rehear Perplexity agentic AI vacated injunction rulingMealey's, September 2026
  52. Securing agentic commerceCloudflare, October 2025
  53. Yes, August 2 still mattersJones Walker, 2026
  54. Recent UK legal and regulatory developments on AI and automated decision-makingKennedys, 2026
  55. Revised guidance on model risk management (SR 26-2 attachment)Federal Reserve, OCC and FDIC, April 2026
  56. CAISI issues request for information about securing AI agent systemsNIST, January 2026

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 6, 2026. No client data appears in our insights.

Read next

All insights
  • A chatbot on its own small island answers on a screen, and its only lane ends at a red stop at the island's edge; on the main plinth an AI agent tower cancels the order, queues the refund, notifies the customer and logs each step, every system ticked.

    AI agents Explainer

    AI Agents vs Chatbots vs Agentic AI: What Actually Differs

    For CTOs and business owners deciding whether a workflow needs an AI agent, a chatbot or plain automation before they fund the build.

    16 min read

  • An AI agent on its own tower reads web pages, email and tool text from an island outside; every action it plans passes through a lit policy engine, which lets calls to the company's systems through and stops a call to send data out at a lowered barrier. A kill switch is wired to the agent at the front.

    Security, risk and compliance Guide

    AI Agent Security: The Threat Model and Controls a CISO Should Require

    For CISOs and security architects deciding whether an AI agent is safe to connect to production systems, company data and customers.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.