AI agents Guide
Computer-Use Agents vs RPA in 2026: What AI That Operates Screens Can Replace
AI agents can now read a screen, move the mouse and type into almost any application, which makes them look like a replacement for every RPA bot a company runs. The best of them score above 85% on short desktop benchmarks. On long workflows, repeated runs and real business software, the numbers fall away fast. This guide sets out what the evidence says agents can take over, what bots still do better, and how to route the work between them.
For COOs, CIOs, heads of shared services and automation leads deciding whether to replace, extend or keep their RPA estate as computer-use AI agents arrive.
The short answer
Computer-use agents can take over screen work that RPA never paid for: low-volume tasks, screens that change often and systems without an API. They cannot yet replace bots on stable, high-volume processes, where a bot is faster, cheaper per run and gives the same result every time. Route each step to an API, a bot, an agent or a person, and check every agent write against the system of record.
Key takeaways
- Short benchmarks are nearly saturated, with self-reported OSWorld-Verified scores up to 86.1%, while the best published strict score on the long workflows of OSWorld 2.0 is 48.7%.12
- On real ERP software one agent saved records in up to 85% of runs but wrote the correct value in as few as 3%, an error a screenshot does not show.3
- Repetition is the weak point: a strong agent solved about 78% of OSWorld tasks at least once in ten tries and all ten times on only about 36%.4
- In the one controlled test against UiPath, the bot ran 10 out of 10 times on every task and the agent was slower and less reliable, but took minutes to build where the bot took hours.5
- Microsoft charges 5 Copilot Credits for a computer-use step and 13 credits per 100 deterministic flow actions, about 38 times more per step.6
- Prompt injection is down to fractions of a percent in vendor tests, and OpenAI and the UK NCSC both say it may never be fully solved.278
A computer-use agent is an AI model that operates software the way a person does. It takes a screenshot, decides what to click or type, acts, and looks again, until the task is done. Robotic process automation (RPA) also works through the screen, but a bot follows steps a developer recorded in advance and breaks when a field moves. The agent decides its path at run time, so it copes with changing layouts and messy input, and it can also choose a wrong path. Since 2025 every major model vendor and every RPA vendor has shipped a version of computer use, and the question for anyone running a bot estate is which work moves, which stays, and how to prove the agent did it right.
- 36%of OSWorld tasks a strong agent completed on all ten of ten runs, against about 78% completed at least once4
- 3%correct values written by some agents on ERP tasks where they saved the record in up to 85% of runs3
- 38xthe cost of a computer-use step against a deterministic flow action in Microsoft Copilot Studio6
What the benchmarks measure and what they leave out
The best-known test is OSWorld, a set of desktop tasks in a virtual machine. On its cleaned-up version, OSWorld-Verified, the top 2026 entries run from 83.4% to 86.1%, and every one is reported by the vendor that built the model. The leaderboard that collects them warns that rows "can vary by evaluator, harness, attempt budget, tool access, task filtering, or verification level".1 The human score usually set beside them, 72.36%, was measured on the original 2024 task set.9 Stanford's AI Index 2026 put agents at 66.3% on OSWorld, "within 6 percentage points of human performance", in data that is already six months old.10
Longer work tells a different story. OSWorld 2.0, released in June 2026, has 108 workflows that take a person a median of about 1.6 hours and an agent an average of 318 tool calls, against about 30 in the original. At launch the best agent completed 20.6% of them.11 Anthropic's system card for Claude Opus 5.5 reports 48.7% completed on a strict measure and 81.8% on partial credit.2 Tests built on real business applications show the same drop as tasks get longer.
| Benchmark | What it tests | Best result | Human reference |
|---|---|---|---|
| ERPBench (Accenture, Sept 2026) | 30 tasks in the open-source ERPNext, scored on database values | Claude Sonnet 4.6 completed 94% of single-record tasks and 100% of the two harder tiers; the strongest open models reached 34% and 32% on single records3 | 87% to 100% across three annotators |
| LegacyWorld (Aug 2026) | 28 legacy Windows business and healthcare workflows | Claude Opus 4.6 78.6% valid success; GPT-5.4 3.6%14 | Not published |
| UI-CUBE (UiPath, Nov 2025) | 226 enterprise UI tasks | 67% to 85% on simple actions, 9% to 19% on complex workflows12 | 97.9% simple; 61.2% complex for people new to the apps |
| SaaS-Bench (May 2026) | 106 tasks across 23 open-source business applications | 3.8% fully resolved by Claude Opus 4.713 | Not published |
Two lessons follow for buyers. A high general score does not predict results on your software: in ERPBench, Holo3-35B-A3B and Qwen3-VL-32B completed 34% and 32% of the simplest ERP tasks while Claude completed 94%.3 And the drop comes with length. UiPath's own researchers describe "a sharp capability cliff rather than gradual performance degradation" between simple and complex tasks.12 We found no independent public test of computer-use agents on SAP, Workday or mainframe screens, so any claim about those systems needs your own trial.
Why one successful run is the wrong measure
A bot that posts invoices runs the same process thousands of times a month, so what matters is whether it succeeds every time. A 2026 study ran a strong agent, Agent S3 with GPT-5, ten times on each OSWorld task. It completed about 78% of tasks at least once, and all ten times for only about 36%.4 SaaS-Bench saw one model score anywhere from 0 to 0.679 on the same task across runs, and set out the arithmetic of long work: at 95% accuracy per checkpoint across 12 checkpoints, the whole task succeeds only about 54% of the time.13
Most failures are failures of judgment. An audit of failed runs across five benchmarks traced 35.2% of genuine failures to planning and 39.3% to verification and feedback, with the largest single cause, at 29.5%, an agent repeating an action that did nothing. Clicking the wrong thing accounted for 13.9%.15 The failure that matters most to a finance team is the silent one. In ERPBench some agents saved the record in up to 85% of runs and wrote the correct value in as few as 3%.3 In LegacyWorld, Claude Opus 4.6 left unwanted changes behind in 10.7% of runs and Kimi K2.5 in 35.7%.14 SaaS-Bench documents agents claiming success after their own verification failed.13
Benchmarks carry their own error. The same audit found 15.3% of failure verdicts were wrong, 17.5% on OSWorld and 21.7% on WebArena.15 Agents are also slow. The best take 2.7 to 4.3 times more steps than necessary, and a late step can take three times longer than an early one.16 In ERPBench, Claude Sonnet 4.6 spent 89 seconds and 231,000 input tokens on a single record, and 323 seconds and about 2 million input tokens on a chained workflow.3 At today's $2 per million input tokens for Claude Sonnet 5.5, that is roughly $0.46 and $4 a run before caching, by our arithmetic.17
The one controlled test against an RPA bot
Researchers at the Technical University of Liberec built the same three processes in UiPath and with Anthropic's computer-use agent on Claude Sonnet 4, then ran each ten times. The tasks came from rpachallenge.com, a set of exercises the RPA community uses to test bots.5
| Task | UiPath bot | Computer-use agent | Time to build |
|---|---|---|---|
| Copy spreadsheet rows into a web form whose layout changes each round | 139.8 seconds, 10 of 10 runs | Did not complete its one run | About 40 minutes for the bot; the agent never reached a working version |
| Watch a stock price and alert below a threshold | 53.9 seconds, 10 of 10 | 109.8 seconds, 9 of 10 | About 38 minutes against about 10 |
| Read invoices and enter their data | 20 seconds, 10 of 10 | 202.8 seconds, 6 of 10 | About 240 minutes against about 15 |
The authors call the agent "not yet production-ready" and found it far quicker to build.5 The bot was faster on every task, and the differences in reliability were too small a sample to be statistically significant. The study used a 2025 model, and later benchmarks point the same way: agents trade run time and repeatability for build time and tolerance of change. That trade decides where each one belongs.
What computer use costs
Model vendors now bill computer use as ordinary tokens, and the buyer runs the browser or virtual machine. Microsoft and Amazon sell it as a metered unit. RPA vendors sell bots by the month and agent calls by platform credit.
| Product | Status, October 2026 | Published price | Where it runs |
|---|---|---|---|
| OpenAI computer use in the Responses API | Current, replacing computer-use-preview18 | GPT-6 Astra $10 input and $50 output per million tokens; GPT-6.1 Sol $2 and $1019 | Your browser or virtual machine |
| Anthropic computer use tool | Claude API and Google Cloud; beta on Amazon Bedrock, Microsoft Foundry and Claude Platform on AWS20 | Claude Sonnet 5.5 $2 and $10; Opus 5.5 $4 and $20; the tool adds about 4,500 input tokens per request17 | Your virtual machine or container |
| Google Gemini API computer use | Gemini 3.5 to 3.8 Flash models21 | Gemini 3.8 Flash $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.5022 | Your client environment |
| Microsoft Copilot Studio computer use | OpenAI's agent and Claude Sonnet 4.5 generally available; Claude 4.6 models experimental23 | 5 Copilot Credits a step, 15 on premium models; $200 for 25,000 credits, about $0.04 or $0.12 a step; not included in Microsoft 365 Copilot user licenses624 | A Windows machine |
| Amazon Nova Act | Available on AWS | $4.75 per agent hour of elapsed working time25 | AWS |
| Microsoft Power Automate (RPA) | Current | $15 per user a month; $150 per unattended bot a month; $215 per hosted bot a month26 | Your Windows machines or Microsoft-hosted |
| UiPath | Current | Basic from $25 a month, higher tiers by quote; each agent model call costs 0.16 to 0.4 Platform Units, counted per 64,000 input tokens2728 | UiPath cloud or your own servers |
Microsoft is the one vendor that prices a bot, a deterministic flow and a computer-use step side by side, which makes it a clean comparison. A flow action costs 13 credits per 100 actions and a computer-use step costs 5 credits, about 38 times more.6 One hosted Power Automate bot at $215 a month buys the same as about 26,900 credits, or roughly 5,400 standard computer-use steps.2624 At 30 steps a task, that is about 180 agent tasks a month. The arithmetic leaves out build and maintenance labor, which is most of what RPA costs, but it shows where the line falls. On a stable screen running several hundred times a month, the bot is cheaper to run. On a screen that runs a few dozen times a month or changes every quarter, the agent's shorter build and lower upkeep can win.
How RPA has actually performed
The most quoted statistic about RPA, that 30% to 50% of projects fail, comes from a 2016 EY paper, where it is an observation about first projects: "we have seen as many as 30 to 50% of initial RPA projects fail".29 There is no survey behind it. The better evidence points at upkeep, scale and measurement. HFS Research estimated that licenses are "just 25% to 30% of total costs for implementing RPA".30 Deloitte's 2020 survey found 37% of organizations piloting with 1 to 10 automations and 13% scaling beyond 50.31 Its next survey found the average payback period for those still piloting had grown from 16 months in 2020 to 22 months.32
US government auditors give the clearest record. The inspector general for the General Services Administration found that the agency's claim that its RPA program reclaimed more than 240,000 work hours a year was "inaccurate and unreliable", and that GSA was not tracking what its bots cost.33 A 2024 follow-up counted 119 active and 24 decommissioned bots, found no process for removing retired bots' access, and found that program management "simply removed or modified the requirements" it could not meet.34 HUD's inspector general concluded that after more than three years its program "had achieved minimal progress and results".35 Measurement, cost tracking and machine identities are the same three problems agents bring, with less predictable behavior on top.
The RPA business itself is growing slowly. UiPath reported $1.938 billion of annual recurring revenue at July 31, 2026, up 12%, with $37 million net new.36 Eighteen of its top 20 deals that quarter included AI, and its chief executive described the pitch plainly: AI "is probabilistic and can be expensive at scale", while many processes "need exactness, the same result every time".37 Automation Anywhere says AI accounts for nearly 70% of its new and upsell bookings and that agent executions grew five times in a year.38 Analysts expect slow adoption: Forrester predicted that fewer than 15% of firms will turn on the agentic features in their automation suites in 2026,39 and Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027.40
What deployments look like so far
We found no company that has publicly retired a set of RPA bots in favor of agents and published the before and after costs. The public cases follow two patterns. The first is computer use where no API exists. Graebel, a relocation company, uses Copilot Studio computer use because its globalCONNECT platform "doesn't expose APIs for this workflow". The agent reads service orders, works through the platform's screens and routes "low-confidence cases, exceptions, and approvals for human review". Microsoft describes it as a pilot covering "a small set of high-volume service order types" and publishes no figures.41
The second pattern is agents working alongside existing bots, which is how the RPA vendors sell. UiPath told investors it handles about 700,000 invoices a year for a Fortune Global 500 manufacturer, and reported 96% document accuracy in a proof of concept.37 Amazon says Nova Act reached 90% reliability on browser workflows built by early customers, and that one startup client automates hundreds of thousands of workflows a month with it.42 These are vendor figures without independent measurement. In outsourced back-office work the effect so far shows up in revenue mix: Genpact's core business services grew 1.9% in its second quarter of 2026 while its technology business grew 24.1%.43
Prompt injection is lower in vendor tests and still open
A computer-use agent reads untrusted content, web pages, emails, documents and screenshots, and acts on it with the access it was given. Text planted in that content can redirect it. Vendor measurements have fallen fast. In August 2025 Anthropic measured a 23.6% attack success rate for browser use without its mitigations and 11.2% with them.44 A year later, on a harder set of attacks written by professional red-teamers, attacks succeeded against Claude Opus 4.5 17.6% of the time and against Opus 5 3.8% of the time before extra safeguards.45 The Claude Opus 5.5 system card reports 0.07% in computer-use environments, two successful attempts out of 2,800. The same card reports 54.61% for an adaptive attacker in coding environments, falling to 11.13% with probes enabled, which shows how much the number depends on the setting.2
No one who builds these systems calls the problem solved. OpenAI has said its AI browser may always be vulnerable to prompt injection.7 The UK's National Cyber Security Centre said such attacks "may never be totally mitigated" the way SQL injection can be, and told developers to focus on secure design.8 Independent researchers keep finding working attacks. Brave showed in August 2025 that a hidden comment on Reddit could make Perplexity's Comet browser log into the user's account, read a one-time password from Gmail and send both to the attacker.46 In academic tests, visual prompt injections fooled computer-use agents up to 51% of the time and browser agents up to 100% on some platforms,47 and WASP found attacks partly succeeding in up to 86% of cases while agents completed the attacker's goal in 0 to 17%, which its authors call "security by incompetence".48 That margin shrinks as agents get better at finishing tasks.
| Risk | Control | What an auditor can check |
|---|---|---|
| Instructions planted in pages, emails or screenshots | Isolated virtual machine, domain allowlist, no standing credentials in the session | The machine image, the allowlist and the credential vault's access log |
| A wrong value saved to a system of record | A check against the database or API after every write, with mismatches sent to a queue | Check results per run and the mismatch queue |
| Payments, purchases, deletions and messages sent outside | A person confirms before the action | Approval records with the approver's identity |
| Access wider than the task | A separate least-privilege identity for each agent, reviewed and removed on retirement | The identity inventory and review dates |
| Behavior nobody can reconstruct | Screenshots and actions recorded for every run and kept | Session logs tied to transaction numbers |
Legal and regulatory limits
| Rule or case | Status, October 2026 | What it means for computer-use agents |
|---|---|---|
| Amazon v. Perplexity, US Court of Appeals for the Ninth Circuit | Injunction against Perplexity's Comet agent vacated on August 4, 2026; rehearing by the full court declined5051 | Under US anti-hacking law it is the user who "accesses" a site through an agent. Contract, unfair competition and other claims remain open, so a portal's terms bind the company directing the agent |
| Web Bot Auth for signed agents | Used by Visa, Mastercard and American Express for agent commerce through Cloudflare52 | Sites are starting to admit agents that prove who they are. Unsigned screen automation will meet more blocking on third-party sites |
| EU AI Act Article 50 | Most transparency duties apply from August 2, 2026; Annex III high-risk duties moved to December 2, 202753 | An agent that writes to people, such as suppliers or customers, must not pass as a person. Agents in hiring or credit decisions fall under the later high-risk regime |
| UK GDPR Articles 22A to 22D | In force since February 5, 202654 | Solely automated decisions with significant effects are allowed more widely in the UK, with safeguards. The EU rules on such decisions remain stricter |
| US bank model risk guidance, SR 26-2 | Replaced SR 11-7 on April 17, 202655 | Generative and agentic AI "are not within the scope of this guidance", so banks must set their own controls for agents |
Which mechanism for which step
The decision is made step by step, and one process usually ends up mixing an API call, a bot, an agent and a person. The order of preference that the evidence supports is an API first, a deterministic bot second, an agent third, and a person for anything that cannot be checked automatically.
| Situation | Use | Why |
|---|---|---|
| The system has a documented API or connector | API or integration platform | Deterministic, cheapest per run and fully logged |
| Stable screen, hundreds of runs a month or more, structured input, no API | RPA bot | Faster and repeatable; at volume the bot's monthly fee is a fraction of the agent's steps |
| Low volume, many different screens, layouts that change, no API | Computer-use agent with a check after every write | Minutes to build and tolerant of layout changes |
| Emails, PDFs or free text feeding a structured system | A model reads and classifies; an API or bot writes | Keeps the uncertain step away from the write |
| Exceptions from an existing bot | Agent on the exception queue, bot on the main path | Judgment where it adds value, exactness where the volume is |
| Long workflow across several applications | Split into stages with a state check between each | End-to-end success on SaaS-Bench was 3.8% |
| SAP, mainframe terminals, Workday, Oracle | Keep the API or bot; trial an agent beside it | No independent public results on these systems |
| Decisions with legal or similar effects on a person | A person decides, the agent prepares | GDPR and the UK safeguards require meaningful human involvement |
API or integration
When the system offers one
- Same result every time
- Cheapest per run
- Survives screen redesigns
- Needs the vendor to expose the operation
RPA bot
Stable screens at volume
- Runs in seconds
- Predictable cost at high volume
- Breaks when fields move
- Hours to build and to repair
Computer-use agent
The long tail and changing screens
- Minutes to set up
- Copes with layout changes and messy input
- Slower and varies between runs
- Needs a check after every write
How to start in eight weeks
- List the screen work Inventory every bot and every manual screen task with its monthly volume, how often the screens change, and whether the system has an API.
- Pick the long tail Choose three to five low-volume tasks with no API and frequent screen changes, which bots never covered economically.
- Define correct For each task, write the database or API check that proves the work was done right, independent of what the agent reports.
- Build the sandbox Give the agent its own virtual machine, its own least-privilege identity, a domain allowlist and no stored passwords.
- Run it ten times Run every task repeatedly on real cases and count how often it succeeds on all runs, alongside the time, steps and cost per run.
- Price at real volume Compare the cost per task at your monthly volume, including review time for exceptions, with a bot or a person doing the same work.
- Go live with a queue Send low-confidence cases and failed checks to people, record every session, and widen the scope only when the numbers hold.
Questions before replacing a bot with an agent
- Does this system have an API we could use instead of its screens?
- How many times a month does this task run, and how often do its screens change?
- How will we prove each write was correct without trusting the agent's own report?
- How often does the agent succeed on every one of ten runs on our real cases?
- What does a run cost at our volume, including people reviewing exceptions?
- What can the agent's identity reach if planted text redirects it?
- Do the terms of any third-party portal it uses allow automated access?
Clicking is close to solved in 2026. Doing a long task correctly every time, and proving it, is not, and that is the thing RPA's rigid design was built to guarantee. Agents earn their place where bots were always too expensive to build and maintain: the low-volume screens, the changing layouts and the messy input that stayed manual. The organizations that gain most will be the ones that write the checks, measure repeated runs and route each step to the cheapest mechanism that can be trusted with it.
This is how we approach screen automation in our AI agent development work: inventory the bot estate and the manual screen work, use APIs wherever they exist, give agents the long tail inside an isolated sandbox with their own identity, and put a database check and an exception queue behind every write.
Questions leaders ask
Will AI agents replace RPA?
Not on stable, high-volume processes in 2026. In the only controlled comparison, a UiPath bot was faster and succeeded on every run, while the computer-use agent was slower and less reliable. Agents extend automation to low-volume screens, changing layouts and unstructured input that bots never covered economically, and they increasingly work alongside existing bots.
What is a computer-use agent?
An AI model that operates software through its screen. It takes screenshots, decides what to click or type, and repeats until the task is done. OpenAI, Anthropic, Google, Microsoft and Amazon all offer one, and RPA vendors such as UiPath and Automation Anywhere have added agents to their platforms.
How reliable are computer-use agents in 2026?
Good on short tasks and weak on long ones. The best agents score above 80% on short desktop benchmarks and 94% on single ERP records, but only 48.7% on 1.6-hour workflows and 3.8% on tasks across several business applications. Repeated runs matter most: one strong agent solved all ten runs on only about 36% of OSWorld tasks.
How much does a computer-use agent cost per task?
It depends on the steps and the model. Microsoft charges about $0.04 a step on standard models in Copilot Studio, Amazon charges $4.75 per agent hour for Nova Act, and model vendors bill tokens. A chained ERP task that used about 2 million input tokens would cost roughly $4 at $2 per million before caching.
Is it safe to give a computer-use agent company credentials?
Only with the access limited to what the task needs. Run the agent in its own virtual machine with its own least-privilege identity, a domain allowlist and no stored passwords, and require a person to confirm payments, deletions and outside messages. Vendors have cut prompt injection rates sharply, and none of them says the problem is solved.
Should we use computer use when the system has an API?
No. An API call gives the same result every time, costs far less per run and survives screen redesigns. Computer use is for systems without an API, or where the API does not expose the operation you need.
Can computer-use agents automate SAP or mainframe screens?
There are no independent public results for SAP GUI, Workday, Oracle E-Business Suite or mainframe terminals as of October 2026. If an API or a working bot exists, keep it, and trial an agent beside it on your own cases, with a database check behind every write.
Sources
- OSWorld-Verified leaderboardSteel.dev, 2026
- Claude Opus 5.5 System CardAnthropic, 2026
- ERPBenchAccenture, arXiv 2609.17885, September 2026
- On the reliability of computer use agentsarXiv 2604.17849, April 2026
- RPA and LLM agents compared on enterprise tasksPrůcha, Matoušková and Strnad, arXiv 2509.04198, September 2025
- Billing rates and managementMicrosoft Learn
- OpenAI says AI browsers may always be vulnerable to prompt injection attacksTechCrunch, December 2025
- Mistaking AI vulnerability could lead to large-scale breachesUK NCSC, December 2025
- OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environmentsXLANG Lab
- AI Index Report 2026: technical performanceStanford HAI, 2026
- OSWorld 2.0arXiv 2606.29537, June 2026
- UI-CUBEUiPath, arXiv 2511.17131, November 2025
- SaaS-BencharXiv 2605.15777, May 2026
- LegacyWorldarXiv 2608.14131, August 2026
- How benchmarks mis-score computer-use agentsarXiv 2607.28367, July 2026
- OSWorld-Human: benchmarking the efficiency of computer-use agentsarXiv 2506.16042, June 2025
- Claude API pricingAnthropic
- Computer useOpenAI API docs
- API pricingOpenAI
- Computer use toolAnthropic docs
- Computer useGoogle Gemini API docs
- Gemini Developer API pricingGoogle
- Automate web and desktop apps with computer useMicrosoft Learn
- Copilot Studio pricingMicrosoft
- Amazon Nova pricingAWS
- Power Automate pricingMicrosoft
- PricingUiPath
- Agents licensingUiPath docs
- Get ready for robotsEY, 2016
- How to make sense of nonsensical RPA software pricingHFS Research, 2018
- Automation with intelligence: 2020 surveyDeloitte, 2020
- Intelligent automation 2022 survey resultsDeloitte, 2022
- GSA's robotic process automation program lacks evidence to support claimed savingsGSA Office of Inspector General, November 2023
- GSA's robotic process automation program: security auditGSA Office of Inspector General, August 2024
- HUD's robotic process automation program was not efficient or effectiveHUD Office of Inspector General, 2023
- UiPath reports second quarter fiscal 2027 financial resultsUiPath, September 2026
- UiPath Q2 2027 earnings call transcriptThe Motley Fool, September 2026
- Automation Anywhere reports continued double-digit growth in Q2 FY27Automation Anywhere, September 2026
- Predictions 2026: automation at the crossroadsForrester, 2025
- Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027Gartner, June 2025
- Graebel agent transformation storyMicrosoft
- AWS agentic AI announcementsAbout Amazon, December 2025
- Genpact reports second quarter 2026 resultsGenpact, August 2026
- Piloting Claude for ChromeAnthropic, August 2025
- Claude in Chrome is generally availableAnthropic, August 2026
- Agentic browser security: indirect prompt injection in Perplexity CometBrave, August 2025
- VPI-Bench: visual prompt injection attacks for computer-use agentsarXiv 2506.02456, 2025
- WASP: benchmarking web agent security against prompt injection attacksarXiv 2504.18575, 2025
- OWASP Top 10 for Agentic ApplicationsOWASP GenAI Security Project, December 2025
- Ninth Circuit holds human user, not developer of AI agent, responsible for websites accessedTroutman Pepper Locke, August 2026
- 9th Circuit won't rehear Perplexity agentic AI vacated injunction rulingMealey's, September 2026
- Securing agentic commerceCloudflare, October 2025
- Yes, August 2 still mattersJones Walker, 2026
- Recent UK legal and regulatory developments on AI and automated decision-makingKennedys, 2026
- Revised guidance on model risk management (SR 26-2 attachment)Federal Reserve, OCC and FDIC, April 2026
- CAISI issues request for information about securing AI agent systemsNIST, January 2026
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 6, 2026. No client data appears in our insights.