Abstract digital data landscape representing email security and threat protection

The Double Agent Problem: Why Your AI Agents Can't Be Trusted by Default

Share with your network!

In espionage, a double agent appears to serve one side while working for another. What makes a double agent dangerous is legitimate access: they hold the clearance, sit in the briefings, and handle the documents, so every record shows an authorized person doing authorized work. AI agents create this condition by default. 

When you connect an agent to your email, CRM, cloud storage, and internal tools, you grant access to a reasoning system that chooses its actions as it goes, to accomplish its task or goal. A conventional application does exactly what its code specifies, so a script written to read one directory and send a summary email does that every time. An agent given the same access interprets your request, decides which steps will accomplish it, and executes those steps with the tools available to it. 

That ability is why organizations deploy agents, and it is also where the risk begins. An agent stays aligned with your intent only as long as it interprets your request correctly at each step. That interpretation can fail in three ways: the agent can get it wrong, lose it partway through a workflow, or have it overridden by indirect prompt injection, where an attacker hides instructions in content the agent processes.

The first two failures need no attacker, and in all three the agent keeps working with legitimate access while its activity looks normal in the logs. 
 

AI Agents Exceed the Scope of Their Tasks without Any Attacker Involved

In February 2026, the director of alignment at Meta Superintelligence Labs lost control of an OpenClaw agent connected to her email. She had asked the agent to review her inbox and suggest what to delete or archive, and she told it to confirm before acting. Her inbox was large enough to trigger context window compaction, and the compressed history the agent kept dropped her instruction to confirm. The agent then treated its job as cleaning the inbox, deleting more than 200 emails and ignoring the stop commands she sent from her phone until she ran to the machine hosting it and killed the process. 

A month later, an internal agent at Meta caused an incident the company rated Sev 1, its second-highest severity level. An engineer asked the agent to analyze a technical question a colleague had posted on a company forum, and the agent posted its own answer to the forum without the engineer's approval. The answer was wrong, and when the colleague followed it, sensitive company and user data became visible to unauthorized engineers for about two hours. 

Neither incident involved an attacker: each agent used access it legitimately held to do something its user never asked for, a failure known as semantic privilege escalation. 

Classic privilege escalation happens when an attacker exploits a vulnerability to gain access beyond what they are authorized to have, such as moving from a standard account to administrator. Semantic privilege escalation happens when an agent uses its authorized permissions to take actions beyond the scope of the task it was given. The permissions are valid, but their use is inappropriate given the context. 

The controls security teams rely on are not built to catch this failure, because they treat trust like a clearance, granted once and held until revoked. Access control asks whether an identity may perform an action, and both agents had access for everything they did. Insider threat programs look for activity that departs from a behavioral baseline, and an agent can drift from the user's intent after a single context window compaction or one iteration of its reasoning loop while its activity still looks like the work it was deployed to do. 
 

Attackers Hijack Agents through Indirect Prompt Injection 

An attacker can use indirect prompt injection to cause the same failure on purpose. The attacker needs no credentials and no foothold in the network, only a way to put text in front of the agent, and Salesforce Agentforce shows how little that takes. Agentforce agents can process leads submitted through public web-to-lead forms, which accept text from anyone on the internet, and researchers have repeatedly shown that an agent will follow instructions hidden in those form fields. 

Researchers first reported the technique in September 2025, showing that hidden instructions in a lead form could make an agent send CRM data to an attacker-controlled domain. Salesforce responded by enforcing a list of trusted URLs for its AI agents. In April 2026, another research team showed that a single injected line in a lead form could still get an agent to collect every lead it could find and email the list to the attacker. 

In September 2026, a third team bypassed Salesforce's URL controls and went further. Through a malicious lead, the researchers made an Agentforce agent post a reply in an internal Slack thread. The reply could carry a phishing link and gave no sign that an agent wrote it, so employees had every reason to trust it as a message from a colleague or the IT help desk. Salesforce has since changed its defaults to require user confirmation before agents send Slack messages, and it says it has seen no evidence of exploitation. 

In all three cases, the attacker never accessed the organization's systems directly. Like a handler who never enters the building, the attacker relied on the hijacked agent to carry out the attack with the access it had been granted to serve legitimate users. 

Neither of the obvious responses solves this: Salesforce's URL controls held only until researchers disguised their exfiltration domains in formats the controls did not recognize, and no filter can anticipate every way an attacker might phrase an instruction. Revoking the agent's email and Slack access would stop the attack, and it would also stop the agent from doing the sales work it was deployed for. 

Counterintelligence catches double agents by checking what an operative does against the mission they were given. AI agents need the same check on every action at runtime, and the check has to be automated, because a single request can trigger dozens of actions, far more than a security team can review by hand. 
 

Proofpoint AI Security Verifies Every Agent Action at Runtime 

Proofpoint AI Security checks each action against three things: what the user asked for, what the agent was deployed to do, and what the organization allows. Proofpoint's Intent-Based Access Control (IBAC) handles the first, agent manifests handle the second, and Semantic Business Policies handle the third. 

IBAC captures the intent of a request when the workflow begins, as a semantic understanding of what the user is trying to accomplish. At runtime, a purpose-built model evaluates each tool call, data access, and LLM interaction against that intent, weighing the type of action, the data involved, the sequence of earlier actions, and the workflow expected for the request. 

In the OpenClaw incident, the user asked only for suggestions, so deleting email fell outside her intent. IBAC detects that deletion at runtime and, depending on policy, blocks it or holds it for her approval. Because IBAC captured her intent when the workflow began, the check still works after context window compaction drops her instruction from the agent's history. 

The same check applies to the Meta incident. The engineer asked the agent to analyze a question, and posting an answer to the forum falls outside that request, so IBAC detects the post and, depending on policy, holds it for the engineer's approval. 

Evaluating actions also avoids the false positives that prompt injection filters produce. A user who asks a financial analysis agent to evaluate a stock and "ignore recent market volatility" is using exactly the kind of instruction-like phrasing those filters flag. Every action the agent takes afterward matches the request, so IBAC lets the work proceed. 

The Salesforce attacks show why Proofpoint evaluates actions against more than the user's request. Agents that process inbound leads work on data anyone on the internet can submit, and their actions are not always tied to a request from a specific user. For those agents, the checks that carry the most weight are what the agent was deployed to do and what the organization allows. 

Proofpoint AI Security defines what an agent was deployed to do with an agent manifest, for first-party agents an organization builds and third-party agents it adopts. Security teams create each manifest in the Proofpoint platform, either by writing it or by approving one Proofpoint generates from the agent's observed behavior. The manifest is policy as code, written in YAML and modeled on Kubernetes deployment specs, so teams can version and audit it like any other infrastructure configuration. 

Each manifest declares the tools, MCP servers, LLMs, and data sources the agent uses, the destinations each of them may communicate with, and the security controls that apply to the agent, including IBAC. Proofpoint enforces the manifest at runtime and denies anything it does not declare by default, so an agent that tries to connect to a destination outside its manifest, such as an attacker-controlled domain, is blocked or flagged depending on policy. 

The Salesforce agent also misused tools its role legitimately includes. An agent that handles leads may need to send email, and one deployed to Slack needs to post messages. What made these actions wrong was where the data went and the missing approval, and those are questions of business policy. 

Proofpoint's Semantic Business Policies let administrators write those rules in plain English, such as "Never send CRM records to an external email address" or "Never let an agent post to an internal channel without end-user approval." Proofpoint generates the runtime controls that enforce each rule, and IBAC evaluates every agent action against them. The same rules govern an employee working through an AI assistant and an agent working on its own. View the self-led demo here.

Organizations decide what happens when an action fails any of these checks. Proofpoint can block the action inline, flag it for human review, or log it for analysis. Organizations typically start in visibility mode, move to detection as their policies mature, and turn on inline enforcement for workflows that involve sensitive data or irreversible operations. 


The Bottom Line

Securing AI agents means verifying every action at runtime against what the user asked for, what the agent was deployed to do, and what the business allows. We've worked closely with customers and industry experts to create the Agent Integrity Framework and Maturity Model, a prescriptive guide for securing AI agents across the enterprise. Download it today.