Data Exfiltration Detection in Agent Tool-Call Sequences
Detecting exfiltration requires watching sequences of agent calls, not individual permissions.

Data exfiltration by an AI agent rarely looks like theft while it's happening. It looks like a series of ordinary tool calls, each one authorized, each one unremarkable on its own. That's the structural mismatch at the center of this problem: the controls enterprises already have were built to catch a single bad moment, and agentic exfiltration doesn't have one.
Traditional data loss prevention tools assume data moves the way a person moves it. DLP watches for that moment of transfer and flags it. That model works fine when a human is the one doing the moving, because people tend to do one risky thing at a time, and that one thing is visible.
An AI agent working through the Model Context Protocol (MCP) doesn't move data in one step. Each of those calls, read on its own, passes as legitimate. Strung together with intent, they add up to a theft that no single log entry reveals.
Static APIs don't have this problem because they only do what they're told, in the way they're told to do it. An agent is different. It decides what to call next based on what it just saw, adjusts its plan mid-task, and moves through chains of tools that shift with context. Every one of those shifts is a new decision point, and every new decision point is a new place for something to go wrong.
The deeper reason this breaks from ordinary API abuse comes down to how an agent picks its next move. MCP tools are selected because a large language model read a plain-English description of what a tool does and decided it fit the task, not through fixed, typed interfaces that a firewall can validate against a schema. That means there's no rigid contract to check calls against, and no request rate to cap, that will catch an agent doing something harmful while staying inside its technical permissions. The harm doesn't live in any single call. It lives in the combination, and a security model aimed at single calls will never see it.
How the exfiltration sequence unfolds across a tool-call chain
A complete exfiltration through MCP tends to follow a pattern with four distinct stages, and understanding that pattern is the starting point for catching it. Each stage is necessary for the theft to succeed, but none of them, taken alone, would set off a conventional alert.
The first stage is enumeration. The agent calls a tool that lists what it can reach, a directory listing, a CRM record query, a file system walk, mapping out the terrain before it decides where to go. That call looks exactly like something a legitimate task would do, because in most cases, it is.
The second stage is targeted access. Read operations are standard, permitted, and expected of an agent doing its job, so this step draws no more attention than the first.
The third stage is staging. Between the read and the send, the agent may reshape the data, combine it with other records, reformat it, or break it into pieces. This step does the most to obscure intent, since it sits in the middle of the chain, away from both the access point and the transfer point, where a reviewer looking for a smoking gun is least likely to be looking.
The fourth stage is outbound transfer. Every arrow is permitted. The danger is in the path, not in any one point along it.
The giveaway, when it appears, is the join between stages two and four: a sensitive read followed immediately by an outbound send, where the values from the read show up in the arguments of the send. Catching that requires watching inputs and outputs together across calls, not reviewing each call in isolation.
None of this needs to happen quickly, either. Any detection method that only looks within a single session misses exactly the kind of attack built to outlast one.
Tool poisoning and prompt injection as the ignition point, not the exfiltration itself
The sequence described above has to start somewhere, and it usually starts inside the tool environment the agent already trusts, not from an outside break-in. That's the detail enterprises tend to miss: the threat begins inside the authorized surface the agent was built to use.
Tool poisoning is the clearest version of this. A tool's description, the plain-English text that tells the agent what the tool does and when to use it, can carry hidden instructions that a person skimming the interface would never notice but that the model reads and follows. The agent pulls up its list of available tools, absorbs those buried directions along with the legitimate ones, and starts carrying out the exfiltration sequence as if it were just another task someone asked it to do.
A malicious MCP server makes this worse because, from the agent's point of view, it looks exactly like any legitimate tool provider. The agent has no built-in way to check whether a server's intentions match its description, so a server built specifically to exfiltrate data can sit in the environment, offer tools that look normal, and quietly direct the agent toward the sequence described earlier. There's a second version of this risk: a rogue tool can return manipulated results that look like normal output, and the agent, trusting what it receives, acts on that corrupted data and spreads the problem further without any new prompt from a user.
This matters for where detection has to focus. Filtering what a user types catches neither a poisoned tool description nor a manipulated tool output, because neither one originates in the user's prompt. By the time anything resembling a suspicious instruction becomes visible at the prompt level, the sequence described in the last section may already be running.
Why traditional DLP, WAF, and API security controls miss agents
Legacy security tools fail here because every one of them was designed around an assumption that doesn't hold for agents: that a human is behind the keyboard, making one identifiable decision to move data, at one identifiable moment.
API gateways, web application firewalls, and schema validation were all built to handle requests that look the same every time a person makes them. An agent doesn't request the same thing the same way twice. It makes decisions based on what it just saw, and each step in its chain depends on the output of the step before it. A gateway built to check a request against a fixed shape has nothing to check an agent's evolving, context-driven call against.
Firewalls, DLP systems, cloud security posture tools, and network egress monitors all watch individual events: one HTTP call, one file write, one outbound connection. None of them can see the reasoning that led to that call, or the string of earlier calls that built up to it. They see the outbound webhook post in isolation. They have no way to connect it back to the sensitive read that happened three steps earlier.
Endpoint detection and response tools run into the same wall from a different angle. They can see that a file was written or a network connection was made, but they can't see why. The prompt that started the task, the reasoning the agent used to decide what to do next, and the chain linking an original instruction to a specific action taken minutes or hours later, none of that is visible to a tool built to watch files and processes.
Rule-based defenses don't solve this by adding more rules. A rule written to flag any outbound file transfer also flags every legitimate use of that same tool. Security teams either drown in false alarms or turn the rule down to the point where it misses the exact sequences it was meant to catch.
Put together, this means an agent can run the full enumerate, access, stage, transfer sequence from start to finish, using tools it's fully permitted to use, and never once cross a line that a standard enterprise security stack is watching for.
How sequence-level behavioral detection works
Catching this requires watching the whole chain, not any single link in it. Detection has to track the agent's prompts, its reasoning steps, every tool call it makes, the arguments and outputs of each one, and treat that whole record as one connected behavior to analyze, rather than a pile of separate events to check off individually.
That starts with telemetry built for agents specifically, not telemetry built to capture outcomes after the fact. One approach along these lines, the ADR Sensor, records the full causal chain, prompts, reasoning steps, tool calls, and the surrounding environment, closing a gap that standard logging never addressed. Every tool invocation gets logged with its inputs, its outputs, how long it took, which parent task triggered it, and where any data it touched ended up going. It's a graph of everything the agent did during a session, built so that patterns spanning several calls at once can actually be searched for and found.
Within that graph, a handful of signals stand out as the clearest markers of a malicious sequence. An enumeration call whose results feed directly into the arguments of a later sensitive read is a pattern legitimate work rarely needs in that exact shape. The clearest signature of all is a sensitive read followed immediately by an outbound send carrying those same values in its arguments, a pattern only visible once inputs and outputs are tied together across separate calls rather than reviewed one at a time. And low-bandwidth exfiltration, small outbound writes repeated through a permitted tool, each one too small to trip any single threshold but adding up to something substantial over time, maps to MITRE ATLAS technique AML.T0086.
Running every one of these checks through full LLM-based reasoning on every single event would cost far too much at the scale enterprises operate at. In production, this architecture held up across more than 10,000 agent sessions a day at Uber, producing zero false positives on the ADR-Bench evaluation set while still catching 67% of attacks, ahead of three other leading detection approaches by a wide margin on F1-score.
Because a sequence can stretch across more than one session, the detection system has to hold behavioral state across sessions too, not reset its memory the moment one session ends. A tool that only looks within a single session has no way to connect a data-discovery phase that happened yesterday to a transfer that happens today.
The authentication and credential gaps that make sequences possible in the first place
None of this sequence would run as smoothly as it does if agents weren't handed far more access than their actual tasks call for. That gap in how credentials get scoped is what makes the enumeration phase so easy to pull off, and current enforcement has no way to close it while the agent is running.
Token delegation across agents is one of the clearest culprits. Token delegation lets one agent act with another's authority without the receiving system verifying that authority, and it appears at nearly every handoff between agents in a multi-agent setup.
Overbroad permissions compound the problem. Only 8.5% of servers listed in the official MCP registry use OAuth, leaving access scoping across MCP deployments thin. That means the enumeration phase described earlier runs into no real technical resistance in most environments.
The spec is moving to address this. Delegation is converging on RFC 8693 OAuth 2.0 Token Exchange, where an agent presents the human's token as the subject and its own credential as the actor, and the authorization server mints a new token that carries both, scoped narrower than what the human originally had.
None of that spec progress comes with a way to force it. OAuth 2.1 is a strong recommendation inside the MCP ecosystem, not something servers are required to implement, and there's still no built-in way to verify who or what is actually connecting to an MCP server.
What enterprise governance must cover
Everything described so far, the sequence, the poisoning, the blind spots in legacy tools, the detection model, the credential gaps, points to one conclusion: governing this risk takes controls built around the tool call itself. That means explicit authorization for each tool an agent can reach, defined scope boundaries, approval gates before certain actions execute, and audit logs that record every invocation in a form nobody can quietly edit after the fact. Model-level guardrails and prompt-level filtering don't reach any of that.
Most enterprise security leaders don't have full visibility into their own AI identities. Most don't enforce access policies for those identities. Agents are operating deep inside the systems that matter most, largely unsupervised at the level that counts.
The real shortfall isn't a missing policy document. Plenty of organizations have written AI principles and put model-level guardrails in place. Agentic AI needs governance that reaches past the model and the prompt down to the level of the tool call itself, since that's where the sequence described throughout this piece actually lives, and policy language alone was never going to reach it.


