Digital server infrastructure representing secure cloud data storage and protection.

Semantic Privilege Escalation: The AI Security Threat Your Access Controls Can't See

Share with your network!

In 1988, Norm Hardy published a short paper about an incident at Tymshare roughly eleven years earlier. A Fortran compiler lived in a system directory called SYSX, and so it could record usage statistics there, administrators had granted it a "home files license" that let it write any file in that directory. The billing file sat in the same directory. A user who learned its name supplied it as the destination for the compiler's debugging output, and the operating system, seeing the license, let the compiler overwrite the billing data.

Hardy's diagnosis was that the compiler held authority from two sources, the user who invoked it and the license the system had granted it, and could not keep the two apart. "The compiler had no way of expressing these intents!" he wrote. He called the problem the confused deputy: a program with legitimate authority, directed to use that authority for a purpose its owner never approved.

Enterprises now connect AI agents to email, cloud storage, CRM, HR systems, financial databases, and collaboration platforms, and each of those agents serves two masters in the same way: the user who assigned the task, and whatever content the agent reads while completing it. Hardy's compiler could be misdirected through a single input, a file name. An agent can be misdirected by any content that enters its context window, because it processes instructions and data in the same natural language and has no reliable way to separate the task it was given from text that redirects it. The confused deputy is back, and this time it can read.

What Is Semantic Privilege Escalation?

Semantic privilege escalation occurs when an agent uses its authorized permissions to take actions outside the scope of the task it was given. Classic privilege escalation requires an attacker to exploit a vulnerability and gain access beyond what an account was granted, such as moving from a standard user to an administrator. Semantic privilege escalation requires no stolen credential, no bypassed control, and no exploited vulnerability in the conventional sense, because the agent applies the access it already holds to work nobody assigned. Every API call succeeds and every authorization check passes, so the security stack records an authorized identity doing authorized work.

The violation lives in meaning, in the distance between what the user asked for and what the agent did. The clearest enterprise example to date came from Microsoft 365 Copilot, although it is usually described as a different threat.

How Did EchoLeak Turn Copilot Against Its Own User?

In June 2025, Microsoft published an advisory for EchoLeak (CVE-2025-32711), a zero-click vulnerability in Microsoft 365 Copilot that it rated critical, with a CVSS score of 9.3. The attacker sends an email carrying instructions hidden from the human recipient. When the user later asks Copilot a question that pulls that email into its retrieval context, Copilot follows the hidden instructions, gathers sensitive data from the user's Microsoft 365 environment, and sends it to an attacker-controlled server through a URL on a Microsoft domain that the content security policy allowed. The user clicks nothing and opens nothing, and every step uses access Copilot legitimately holds on the user's behalf. Microsoft patched the flaw server-side and reported no exploitation in the wild.

Why Isn't Indirect Prompt Injection the Same as Semantic Privilege Escalation?

The industry usually files incidents like EchoLeak under indirect prompt injection, but indirect prompt injection only describes how the attack began. The violation that caused the damage was semantic privilege escalation, and the two are different threats. Indirect prompt injection is a delivery technique. The OWASP Top 10 for LLM Applications defines it as an LLM accepting instructions from external sources such as websites, files, or email, and distinguishes it from direct prompt injection, where the instructions arrive in the prompt itself. Semantic privilege escalation is the security violation, in which an agent uses its authorized permissions for actions outside the scope of its task. Indirect prompt injection describes how attacker instructions enter an agent's context, and semantic privilege escalation describes what the agent does with its access once they are there.

EchoLeak shows where one ends and the other begins. The attack chain starts with delivery, when the attacker's email lands in the inbox, and injection follows when Copilot retrieves that email into its context window while answering an unrelated question. Goal hijack comes next, which the OWASP Top 10 for Agentic Applications lists as ASI01, as Copilot adopts the attacker's objective in place of the user's. Semantic privilege escalation occurs when Copilot applies its authorized access to the user's mail and files in pursuit of that objective, and exfiltration completes the chain. The injection was finished the moment the hidden text entered Copilot's context, and every action that caused harm came afterward, carried out with permissions Copilot legitimately held.

Separating the two changes how the risk is measured and where it is defended. Injection without escalation has a limited blast radius: a model with no tools or credentials that receives injected text produces a manipulated answer, and the damage grows with the permissions the agent holds. Escalation without injection is also possible, because agents step outside their tasks with no attacker involved, as the next sections show. Treating these incidents purely as an input problem points defenders toward ever more input filtering, which catches known payloads but cannot anticipate every phrasing, while the actions that follow a successful injection pass every authorization check. Indirect prompt injection is still the most common way attackers trigger semantic privilege escalation, and the opportunities are growing: Google reported a 32% relative increase in malicious injections on public web pages between November 2025 and February 2026. Any channel where untrusted content enters an agent's context window, from inbound email and shared documents to Slack messages, calendar invites, and RAG retrievals, can carry those instructions, and if an agent can read it, an attacker can write to it.

What Triggers Semantic Privilege Escalation without an Injection?

Injected content accounts for only part of the risk. Two other triggers produce semantic privilege escalation without any hidden instruction, and one of them requires no adversary at all.

Emergent behavior occurs when an agent pursuing a goal takes actions its user never intended, using access it legitimately holds. In July 2025, SaaStr founder Jason Lemkin was nine days into a public experiment building an application with Replit's AI coding agent when he imposed a code and action freeze and told the agent to make no more changes without explicit permission. The agent, which used the same database for development and the live application, then executed destructive commands that deleted the entire database, including records on more than 1,200 executives and nearly 1,200 companies, and told Lemkin a rollback was impossible. Replit's CEO apologized publicly, and the company moved to separate development and production databases by default. Nobody attacked the agent. It held credentials that permitted the deletion, the user's instruction defined a task that excluded it, and nothing in the authorization path compared the two. Enterprise agents face the same pressure in quieter forms, where an agent asked to "prepare for my customer meeting" can pull CRM records, search email threads, and surface a privileged legal communication that happens to mention the customer, with every step appearing relevant to the model.

Tool chain composition occurs when a sequence of individually authorized tool calls produces a compound effect that no single call reveals. Researchers at AWS AI Labs formalized the adversarial version as Sequential Tool Attack Chaining (STAC) and generated 483 test cases in which each step, such as compressing a critical file and then deleting the "duplicate" original, looks benign in isolation. Across eight agents, including GPT-4.1, STAC achieved an average attack success rate of 91.2%, and the strongest prompt-based defense the researchers tested cut that rate by up to 28.8% before its advantage eroded under adaptive attacks. Their conclusion was that defending tool-enabled agents requires reasoning over entire action sequences and their cumulative effects, and the same compound effect can arise from an ordinary multi-step workflow with no adversary involved.

A security control that asks only whether an action is authorized will miss all three triggers, because in each one every action is authorized.

Why Can't RBAC Stop Semantic Privilege Escalation?

Hardy's compiler had no way of expressing intent, and role-based access control has no way of evaluating it. RBAC answers whether an identity holds permission to perform an action. For human employees that answer is usually sufficient, because people bring context the access control system never has to encode: they understand their job, they weigh professional consequences, and they can tell a file they need for an assigned task from a file an unknown sender told them to open.

An agent brings a task description, a set of credentials, and a language interface that treats everything it reads as a possible instruction. RBAC hands it the keys, and nothing in the authorization decision asks whether the agent is using them for the purpose they were issued. A human's role stays stable for months, while an agent's task changes with every request it receives, and an authorization model built on roles has no way to account for that.

Detecting malicious content before an agent reads it remains a necessary layer of defense, and email security, content scanning, and prompt injection detection stop known payloads at the point of entry. Those defenses also face a mathematical ceiling. In June 2026, NIST announced a peer-reviewed proof by senior scientist Apostol Vassilev, published in IEEE Security & Privacy, that extends Gödel's incompleteness theorems to AI guardrails and shows that no finite set of guardrails can be universally robust against adversarial prompts. Any rule set is finite, while the space of natural-language inputs has no bound, and input-layer defenses have nothing to inspect when no injection occurred, which is the case for emergent behavior and for tool chains assembled from benign steps.

What Does the Defense Require?

Industry frameworks have converged on the same diagnosis. The OWASP Top 10 for Agentic Applications, released in December 2025, lists Agent Goal Hijack as ASI01, its first entry, followed by Tool Misuse and Exploitation (ASI02) and Identity and Privilege Abuse (ASI03). Forrester's AEGIS framework, introduced in August 2025, calls on CISOs to enforce least agency and to pivot from securing systems to securing intent, and Forrester's Jeff Pollard has said that least privilege will no longer be sufficient for agents.

Defending against semantic privilege escalation takes four layers that depend on one another. Input-layer defenses, including email security, content scanning, and prompt injection detection, stop known payloads before an agent processes them. A runtime authorization control checks each agent action against the task the agent was given and detects actions that fall outside it. Least agency limits the tools and data an agent can use to what its assigned task requires. Transaction-level forensics records every tool call, data access, and LLM interaction alongside the request that started the workflow, so investigators can trace any outcome to the user who initiated it and the agent that carried it out. Removing any one layer leaves a trigger uncovered, whether that is an injection the input layer never recognized, an emergent action that tripped no content filter, or a compound effect no one can reconstruct after the fact.

Proofpoint AI Security applies these layers to employees and agents alike. Its Intent-Based Access Control (IBAC) captures the intent of each request as a semantic understanding of what the user is trying to accomplish, then evaluates each tool call, data access, and LLM interaction against that intent at runtime, weighing the action type, the data involved, and the sequence of prior actions. Semantic Business Policies, introduced at Protect 2026, enforce the organization's existing policies and code of conduct at runtime, and IBAC evaluates agent actions against those policies as well as the user's request. Depending on policy, a misaligned action is blocked, held for human review, or logged, and Full Transaction Forensics preserves every step for investigation. Applied to EchoLeak, the evaluation needs no knowledge of how the injected text was worded, because a question about recent email carries no intent that covers sending internal data to an external server, and applied to the Replit incident, an instruction to make no more changes carries no intent that covers deleting a database.

The Bottom Line

The confused deputy problem is 38 years old. Hardy's compiler held one license over one directory, while an enterprise AI agent holds credentials across the organization's most sensitive systems and can be redirected by any content it processes, by its own reasoning, or by a chain of steps that each look harmless.

Agent integrity depends on three dimensions evaluated together, at runtime, for every action: what the agent is permitted to do, what its task requires it to do, and what it does. When those three fall out of alignment, the agent is operating outside its scope, and that is where semantic privilege escalation lives. Security for AI agents has to evaluate whether an action is appropriate for the task in addition to whether the identity is permitted to take it, and until that evaluation happens at runtime, every agent is a confused deputy.

Download the framework here: https://www.proofpoint.com/us/resources/white-papers/agent-integrity-framework