Security, risk and compliance Guide

AI Agent Security: The Threat Model and Controls a CISO Should Require

An AI agent reads content an attacker can write and acts with authority someone delegated to it. This guide gives security leaders the threat model for that combination, the control that answers each threat and a twenty-item list to sign against.

For CISOs and security architects deciding whether an AI agent is safe to connect to production systems, company data and customers.

Published
Reviewed
Reading time
16 min

The short answer

AI agent security protects systems in which software holds delegated authority to read data, call tools and act for people. A CISO should require a threat model mapped to the OWASP Top 10 for Agentic Applications, a separate identity with short-lived delegated credentials for each agent, deterministic limits on what injected content can trigger, a vetted tool and MCP supply chain, a tested kill switch and red-team evidence.

Key takeaways

  • An agent's attack surface is everything it reads multiplied by everything it can do. Review shrinks both terms, starting with the principle OWASP calls least agency.1
  • Treat prompt injection as permanent. The UK's National Cyber Security Centre warns it may never be fully mitigated, so the controls that count are deterministic ones outside the model.4
  • Give every agent its own identity, and have it act through delegated tokens that name the user, work for one audience and expire in minutes.56
  • MCP leaves server authorization optional and A2A leaves agent card signing optional. Write both into your policy and enforce them at your gateway.69
  • Sign off on evidence: twenty items, each backed by a test result, a configuration or a drill record.

Every agent rollout reaches the CISO eventually, and the CISO can stop it there. Security reviews were built for software whose behavior is fixed at release. An agent's behavior is decided at run time by a model that reads whatever lands in front of it, and the agent acts with delegated authority. This guide sets out the threat model, the control for each threat and the evidence to see before signing.

  • 13%of the 600 organizations in IBM's 2025 Cost of a Data Breach study reported a breach of an AI model or application2
  • 97%of those breached organizations lacked proper AI access controls2
  • 9.3the vendor's CVSS score for a 2025 flaw that let an attacker pull data out of an enterprise AI assistant3

What changes when software holds delegated authority

An AI agent changes the threat model because it joins three things security teams have always kept apart: authority to act, a decision-maker that can be persuaded by what it reads, and live connections to systems of record. OWASP puts the root problem plainly: agents cannot reliably tell instructions apart from the content they process.1 The UK National Cyber Security Centre (NCSC) warns that prompt injection may never be fully mitigated the way SQL injection can be, because a model draws no inherent line between data and instruction.4

So the attacker model changes. Misusing a conventional application takes a code flaw or a stolen credential; misusing an agent takes only text in front of it, such as an email to a monitored inbox, a ticket comment, a shared document, a web page or a tool description. The agent does the rest with its own permissions, chaining actions with no person in between. OWASP treats a rogue agent as an insider threat amplified by speed and scale, and recommends adding agents to the insider threat program.1

An agent's attack surface is, roughly, everything it reads multiplied by everything it can do, and review is the work of shrinking both. Sometimes that means a smaller design: OWASP's principle of least agency says to give a workflow no autonomy it does not need, since agentic behavior where none is required widens the attack surface and adds no value.1 Tasks that follow fixed rules are easier to secure as deterministic integrations.

The OWASP Top 10 for Agentic Applications, mapped to controls

The OWASP GenAI Security Project published its Top 10 for Agentic Applications on December 9, 2025. It is the reference list of agent threats, and each entry maps to a control a CISO can ask to see working.1 It extends the OWASP LLM Top 10 into multi-step, tool-using systems: goal hijack is the agentic form of prompt injection (LLM01), and privilege abuse the agentic form of excessive agency (LLM06).1

OWASP riskThe attackControl to requireEvidence to ask for
ASI01 Agent Goal HijackInstructions planted in content the agent reads redirect its goalTools narrow after untrusted reads; allowlisted outbound channels; approval for goal changesInjection tests per input channel
ASI02 Tool Misuse and ExploitationA legitimate tool used harmfully, such as a bulk exportPer-tool scopes, rate caps and argument limits, enforced outside the modelTool profiles; refused-call logs
ASI03 Identity and Privilege AbuseReused credentials, or a privileged agent tricked into actingAn identity per agent; delegated tokens that expire in minutes; authority rechecked at executionIdentity register; token settings; traced delegation chains
ASI04 Agentic Supply Chain VulnerabilitiesA poisoned tool, MCP server or peer agent, or instructions in a tool descriptionPrivate registry; hash-pinned versions; descriptions reviewed as codeRegistry allowlist; signed manifests; review dates
ASI05 Unexpected Code ExecutionGenerated code or unsafe deserialization runs commands on your hostsSandboxes with no root and no default network; generation separated from executionSandbox profile; escape tests
ASI06 Memory and Context PoisoningHostile content seeded into memory or the retrieval indexMemory per user and tenant; a source on every entry; expiry; rollbackWrite-validation logs; a tested restore
ASI07 Insecure Inter-Agent CommunicationMessages between agents spoofed, replayed or alteredMutual authentication; signed messages with nonces; pinned protocol versionsMutual TLS settings; replay tests
ASI08 Cascading FailuresOne fault spreads through chained agents into privileged actionsCircuit breakers, quotas and rate limits; a policy engine outside the plannerBlast-radius caps; breaker trip history
ASI09 Human-Agent Trust ExploitationThe agent's fluency talks a person into approving harmApproval screens, written by code, that state the action and its effectsScreen designs; approval and override rates
ASI10 Rogue AgentsAn agent drifts from its declared behavior or is compromisedBehavioral baselines; a declared tool manifest; kill switch with credential revocationAlert rules; kill switch drill records
The ten risks as OWASP names them. Attacks, controls and evidence are our mapping, drawn from OWASP's mitigation guidance.1

Use the table as a worksheet for each agent: every risk gets a control, an owner and evidence, or a written reason it does not apply. An agent with no memory and no peers can exclude ASI06 and ASI07. No agent can exclude ASI01, because every agent reads something.

Give every agent its own identity and short-lived credentials

Every agent should run under its own non-human identity, with credentials that expire in minutes, cover one task and name the person it acts for. OWASP traces privilege abuse to identity systems built for people: an agent without a governed identity sits in an attribution gap where least privilege cannot be enforced.1

Exhibit 1Three ways to give an agent access

Shared service account

The common default

  • One long-lived secret for every user and task
  • Logs never show the person behind a request
  • Holds every permission any task might need

The user's own token

Impersonation

  • The agent is indistinguishable from the user
  • Issued for another audience, and MCP forbids passing it through
  • A hijacked agent holds every right the user has

A delegated, task-scoped token

Token exchange

  • The agent keeps its own identity, and the token names the user
  • Scoped to the task and bound to one audience
  • Expires in minutes and is revocable per agent or task
Only the right-hand pattern answers who did what, for whom and under which authority. Sources: [5], [6]

The standard mechanism is OAuth 2.0 Token Exchange (RFC 8693, January 2020), which distinguishes impersonation, where the agent becomes indistinguishable from the user, from delegation, where it keeps its own identity and acts on the user's behalf. Nested actor claims record each step of a delegation chain, so a trace shows the user, the orchestrating agent and the sub-agent behind every call.5 MCP adds audience binding: clients name the target server when requesting a token, and servers refuse tokens issued for anyone else.6

OWASP's guidance adds three rules.1 Recheck authority when each privileged action executes, since a user's rights can shrink during a long workflow. Never let an agent cache credentials in memory, where a later session can prompt it to reuse them. Keep signing keys away from the agent entirely, with an orchestrator signing on its behalf.

Treat prompt injection as an architecture problem

Prompt injection is managed by design: assume the agent will sometimes follow an attacker's instructions, and make sure doing so achieves nothing of value. The NCSC recommends deterministic safeguards outside the model that constrain what the system can do.4

OWASP cites a case in which one crafted email made an enterprise AI assistant send confidential emails, files and chat logs to an attacker, with no action by the user.1 The flaw, CVE-2025-32711, was published in June 2025; the vendor scored it 9.3 (critical) and the US National Vulnerability Database 7.5 (high).3 Four design rules carry most of the defense:

  • Privileges follow the least trusted input. After reading content from a party, an agent holds no more authority than that party, as the NCSC advises.4 Having read an inbound email, it can draft a reply; it cannot export the customer table.
  • Break the path out. Injection becomes a breach when the agent can send data somewhere. Allowlist outbound destinations, as OWASP's per-tool profiles do, and hold messages to new external recipients for approval.1
  • Keep raw text away from the planner. The component that reads external content returns typed fields, such as a date or an amount, to the component that chooses tools.
  • Put a policy engine between plan and action. A deterministic check outside the model approves or refuses each high-impact call, so a corrupted plan cannot execute itself.1

Injection classifiers catch known payloads and miss novel ones, so they belong in monitoring and never in the sign-off argument. Approval gates for irreversible steps, covered in our guardrails playbook, complete the defense.

Secure the tool and MCP supply chain

Every tool an agent can call, and every MCP server that exposes one, is third-party code with a direct line into the agent's decisions, so it needs the controls of a production dependency plus review of the text it feeds the model. OWASP calls this a live supply chain, since agents load tools and even peer agents at run time, and counts tool descriptions as attack surface because the agent reads them as trusted guidance.1

Both failures have appeared in the wild. OWASP records the first reported malicious MCP server, published to a public package registry under the name of a legitimate email integration, which quietly copied the emails it sent to the attacker.1 Separately, an open-source MCP client component could be made to run operating system commands on the user's machine just by connecting to a malicious server (CVE-2025-6514, scored 9.6, critical).7

For MCP servers and tools, require:

  • A private registry of approved servers and tools, pinned by content hash, where a changed tool description is a new release.1
  • OAuth on every remote server, with tokens bound to that server and refused everywhere else.6
  • Minimal scopes at first connection, step-up for privileged operations and no wildcard scopes.8
  • Local servers started only after the exact command is shown and approved, inside a sandbox.8
  • A bill of materials for each agent, and a kill switch that revokes one tool across every deployment.1

Protect memory and context from poisoning

Memory poisoning is prompt injection that persists: an attacker plants content in what the agent stores or retrieves, and it shapes later sessions, including other users' sessions when stores are shared.1 Treat the retrieval index as memory. If an agent draws on a wiki or ticket history that any employee or customer can edit, every one of those writers can place instructions in front of it. OWASP's controls amount to data governance for everything the agent can remember:1

  • Separate memory per user, per tenant and per domain.
  • Record each entry's source, and weight retrieval by trust.
  • Validate writes before they commit, and expire entries nobody has verified.
  • Never write the agent's own outputs back into trusted memory, where one error becomes self-reinforcing.
  • Keep versioned snapshots so a poisoned store can be rolled back.

Set trust boundaries between agents

Agents should give each other no more trust than an unknown external service gets: every request is authenticated, authorized against the original user's rights and checked against the task. OWASP's counterexample is a crafted email that leads a sorting agent to instruct a finance agent to pay an attacker, and the finance agent complies because the request came from an internal peer.1

The A2A specification, now at version 1.0, requires servers to authenticate every request and scope results to the caller's authority, and it leaves the authorization model to each implementer.9 That model is where your policy lives. Carry the user's identity through the chain with nested actor claims,5 and never give a sub-agent the full rights of its caller. On the wire, require mutual authentication, signed messages with nonces and pinned protocol versions; between planning and execution, add quotas, rate limits and circuit breakers so one fault cannot spread.1

Detection, the kill switch and the incident runbook

Detection for agents compares behavior with a baseline (which tools, in what order, how often, against which data) and alerts on deviation, because each action of a compromised agent can look legitimate on its own.1 It starts with an inventory: in IBM's 2025 study, one in five organizations reported a breach due to shadow AI.2

Exhibit 2Four levels of stop, from narrowest to widest
  1. ToolDisable one tool or MCP server for every agent, with no redeploy.For a suspect component.
  2. AgentPause one agent. New work routes to the human queue.For one agent outside its baseline.
  3. IdentityRevoke the agent's credentials and tokens at the identity provider.For an agent that may ignore a pause.
  4. FleetSuspend autonomous actions everywhere. Reads continue; actions wait for approval.For a compromised shared component.
Each level has a named owner and is drilled before launch. OWASP names kill switches and credential revocation as the response to a rogue agent. Source: [1]

The EU AI Act expects the same of high-risk systems: overseers must be able to interrupt the system with a stop button or similar procedure that halts it in a safe state.12 The runbook then runs in order:

  1. Contain Pull the narrowest stop that works, and rotate every credential the agent could reach.
  2. Preserve Freeze traces, memory snapshots, tool manifests and version records.
  3. Scope Replay the traces to list every record read, action taken and user affected.
  4. Eradicate Purge poisoned memory, roll back tool versions and close the path the attacker used.
  5. Recover Restore service only after fresh attestation of the agent's components and a named person's approval.1
  6. Learn Add the attack to the red-team suite, so the same payload fails on every future release.

Red-team the agent before launch and after every change

Red-teaming an agent means attacking the whole system (inputs, tools, memory and peers) to make it take a harmful action, before launch and after every material change. OWASP recommends periodic tests that simulate goal override and confirm that rollback works.1 Cover each route in:

  • Injection through every channel the agent reads (ASI01).
  • A poisoned tool description or a look-alike server name (ASI04).
  • Planted memory that surfaces in a later session (ASI06).
  • A forged agent card or a replayed message between agents (ASI07).
  • Credential reuse across sessions and users (ASI03).
  • Approval fatigue: harmless requests followed by a harmful one dressed the same way (ASI09).

Score the outcome. A model will sometimes follow an injected instruction; the test passes when the harmful action still fails because a control outside the model stopped it, a measure that holds across model upgrades. After launch, OWASP suggests replaying the previous week's agent actions in an isolated copy of production before any policy is widened.1 A new model, tool or data source is a material change.

Map the controls to NIST AI RMF, ISO/IEC 42001 and the EU AI Act

One body of evidence serves all three frameworks, so build it once. The NIST AI Risk Management Framework, released on January 26, 2023, organizes the work into four functions: govern, map, measure and manage.10 ISO/IEC 42001, published in December 2023, sets requirements for an AI management system.11 Article 15 of the EU AI Act requires high-risk AI systems to resist unauthorized attempts to alter their use, outputs or performance, with measures to prevent, detect, respond to, resolve and control attacks, including inputs designed to make the model err.12 Prompt injection fits that description, so for systems in scope these defenses double as Article 15 evidence.

EvidenceNIST AI RMFISO/IEC 42001EU AI Act (high-risk systems)
Agent inventory with ownersGovern, MapScope and roles of the management systemArticle 9, risk management system
Threat model mapped to the OWASP agentic listMapAI risk assessmentArticles 9 and 15
Identity register and delegated credentialsManageOperational controlsArticle 15
Red-team results across all input channelsMeasurePerformance evaluationArticle 15
Traces with delegation chains and policy versionsMeasureMonitoring and recordsArticle 12, record-keeping
Kill switch drills and incident runbookManageCorrective action and improvementArticle 14, human oversight
Our mapping of the evidence to the three frameworks. Article numbers refer to the requirements for high-risk AI systems in Regulation (EU) 2024/1689.12

The CISO sign-off checklist for an AI agent

Sign when the team shows evidence for all twenty items. An item without evidence stays open as a finding, with an owner and a date.

Design and build: items 1 to 10

  • A threat model covers all ten OWASP agentic risks, each with a control or a reason it does not apply.
  • The agent has its own registered identity and owner; no shared or borrowed credentials.
  • Tokens are delegated, audience-bound and expire in minutes; no long-lived secret is reachable.
  • Authority is rechecked when each privileged action executes.
  • Each tool has a profile: scopes, rate limits, argument limits, destinations.
  • Available tools narrow after the agent reads untrusted content.
  • Outbound channels are allowlisted, and new external recipients need approval.
  • MCP servers and tools come from a private, hash-pinned registry, with descriptions reviewed.
  • Memory is separated per user and tenant, sourced, validated on write and restorable.
  • Generated code runs only in a sandbox without root or default network access.

Launch and operate: items 11 to 20

  • Requests between agents are mutually authenticated, signed and limited to the user's rights.
  • Circuit breakers, quotas and rate limits cap how far a fault can spread.
  • Approval screens use text generated by code, and approval rates are monitored.
  • Traces record inputs, tool calls, delegation chain and policy version for every action.
  • Each agent has a behavioral baseline, and alerts fire on deviation.
  • All four kill switch levels have named owners and have been drilled.
  • The incident runbook has been exercised on an agent scenario.
  • Red-teaming covered every input channel, and no harmful action completed.
  • Model, tool, prompt or data-source changes rerun the red-team suite before release.
  • Evidence is mapped to NIST AI RMF, ISO/IEC 42001 and, where in scope, EU AI Act Article 15.

This is how we approach AI agent development: the threat model comes before the first tool is connected, and this evidence is produced as the system is built, so the security review reads a file instead of commissioning an investigation.

Questions leaders ask

What are the biggest security risks of AI agents?

OWASP's agentic list for 2026 opens with agent goal hijack, tool misuse, and identity and privilege abuse.1 They share a cause: an agent follows instructions found in the content it reads, and it acts with real authority. Poisoned tools and MCP servers, poisoned memory and unauthenticated traffic between agents complete the main routes in. The damage grows with every permission the agent holds.

How do you secure AI agents?

Give each agent its own identity and short-lived delegated credentials, and limit its tools to the task. Assume prompt injection will sometimes succeed, and make it harmless with deterministic controls outside the model. Vet every tool and MCP server like a production dependency, watch behavior against a baseline and keep a tested kill switch. Then prove each control with evidence mapped to the OWASP agentic risks before sign-off.

What is the OWASP Top 10 for Agentic Applications?

It is the OWASP GenAI Security Project's list of the ten highest-impact security risks for autonomous AI agents, published on December 9, 2025.1 The entries run from ASI01 Agent Goal Hijack to ASI10 Rogue Agents and cover tool misuse, identity abuse, supply chain, code execution, memory poisoning, communication between agents, cascading failures and exploitation of human trust. Each entry includes examples and mitigations, which makes it a practical basis for a threat model.

Is MCP secure enough for enterprise use?

MCP can be run securely, and the protocol leaves the key protections to you. Authorization is optional in the specification, so require every remote server to use OAuth, accept only tokens issued for it and never pass tokens through.6 Add a private registry of approved servers pinned by version and review tool descriptions as code,1 and sandbox local servers.8 Without those controls, each server is an unvetted path into the agent.

Should AI agents have their own identities?

Yes. An agent that borrows a person's login or a shared service account cannot be held to least privilege or told apart in the logs, which OWASP describes as an attribution gap.1 Give each agent a registered identity and have it obtain delegated tokens through OAuth token exchange, which records both the agent and the user it acts for.5 Bind each token to one audience and let it expire in minutes.

How is AI agent security different from LLM security?

LLM security is mostly about what a model says: leaked data, harmful text or a manipulated answer. Agent security is about what a system does, because an agent holds credentials and calls tools across many steps. OWASP draws the same line: goal hijack extends prompt injection from one response to a multi-step plan, and privilege abuse extends excessive agency to chains of delegation.1 The controls move from output filters to identity, authorization and containment.

Sources

  1. OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project, December 9, 2025
  2. IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications, 97% of Which Reported Lacking Proper AI Access ControlsIBM, Cost of a Data Breach Report 2025, July 30, 2025
  3. CVE-2025-32711 DetailNational Vulnerability Database, NIST, published June 11, 2025
  4. Prompt injection is not SQL injection (it may be worse)UK National Cyber Security Centre, December 8, 2025
  5. RFC 8693: OAuth 2.0 Token ExchangeInternet Engineering Task Force, January 2020
  6. AuthorizationModel Context Protocol specification, revision 2026-07-28
  7. CVE-2025-6514 DetailNational Vulnerability Database, NIST, published July 9, 2025
  8. Security Best PracticesModel Context Protocol, revision 2026-07-28
  9. Agent2Agent (A2A) Protocol Specification, version 1.0.0A2A Project, The Linux Foundation
  10. AI Risk Management FrameworkNational Institute of Standards and Technology
  11. ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management systemISO and IEC, December 2023
  12. Regulation (EU) 2024/1689, the Artificial Intelligence ActOfficial Journal of the European Union

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.

Read next

All insights
  • An AI agent's road runs into a policy hall that classes every action, and reads and reversible changes run on their own. The irreversible road alone leaves the plinth, across a bridge to the outside world, and a lit gate holds it until a named person beside it approves.

    AI agents Playbook

    AI Agent Guardrails: 8 Controls That Hold Up in Production

    For CTOs and CISOs deciding what an AI agent may do on its own before it touches customers, money or production.

    13 min read

  • On one plinth, your AI agent reaches an ERP, a CRM and a service desk through a single lit MCP gateway; a bridge in the air carries a task over A2A to a partner company's agent on its own island.

    AI agents Guide

    MCP and A2A in the Enterprise: How AI Agents Reach Your Systems

    For CTOs and CISOs deciding how AI agents will connect to SAP, Salesforce and the rest of the estate, and which systems to open to them first.

    16 min read

  • An AI system with its disclosure screen live sends each release down a lane through a release gate, which holds back one failing release, and every release that passes adds a leaf to a tall technical file. A bridge carries the sealed file to a building on its own island, December 2, 2027, when the high-risk obligations begin.

    Security, risk and compliance Analysis

    EU AI Act Compliance in 2026: What the Omnibus Changed and What Is Due Now

    For CTOs, CISOs and general counsel deciding what their AI systems must do to be sold or used in the EU, and what has to be built before December 2027.

    17 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.