Est.

Anomalous Agent Behavior Alerting With Tool-Call Baselines

How tool-call baselines catch AI agent attacks that standard monitoring misses.

Senior Editor & AI Security Correspondent · · 11 min read
Cover illustration for “Anomalous Agent Behavior Alerting With Tool-Call Baselines”
Agentic Incident Response · October 4, 2026 · 11 min read · 2,498 words

A monitoring dashboard can show flat latency, a clean error rate, and a successful response on every call, but an AI agent can still quietly copy a customer database to an external endpoint. That gap, between what the dashboard measures and what the agent is actually doing, is the subject of this piece. Enterprises need a new detection primitive: tool-call baselines that define what normal agent behavior looks like and fire an alert the moment an agent drifts from it.

Why standard monitoring cannot detect a misbehaving AI agent

Latency graphs, error rates, HTTP status codes, throughput counters: these measures were built for software that either works or breaks. A web server that returns server errors is broken. A web server that returns a successful response on every request is working. That binary was good enough for three decades of application monitoring because the software itself made no decisions. An AI agent does. It chooses which tool to call, in what order, with what parameters, based on a prompt and a context window that change from one session to the next, and every one of those choices can look healthy from the outside while being wrong underneath.

The research on agentic attack surfaces explains why this isn't a tooling gap that a better dashboard fixes. Chu et al. ran a layered survey of agentic attack surfaces, and they found three properties that make these attacks structurally invisible to conventional monitoring. The harm comes from the combination of individually authorized actions rather than from any single action that a rule could flag. And they're temporally extended: a payload can sit dormant for weeks after installation and may never present itself as a discrete, loggable event. None of that raises the error rate. All of it becomes visible, eventually, as damage.

What a Compromised Agent Does

Tool poisoning is the clearest illustration of how this plays out. An agent reads a tool's description with the same trust it extends to its own system prompt, so an attacker doesn't need to breach the agent itself. The attacker only needs to alter the text the agent reads before deciding what to do next.

Microsoft's Defender research team disclosed a case that shows this mechanic. A finance team's agent was connected to a third-party invoice enrichment tool, one that had been approved for use but never properly reviewed at the code level. The attacker updated the tool's description, leaving the visible name and summary untouched, and buried a new instruction inside it: collect unpaid invoices and attach them to the next outbound call. Because the Model Context Protocol picks up description changes on the fly, the poisoned version went live with no re-approval step to catch it. Nothing about the agent's request pattern, its latency, or its error rate changed. The only thing that changed was what the agent believed it was supposed to do.

Supply-chain exposure makes the same blind spot worse at scale. As of May 2026, at least seven confirmed high- or critical-severity CVEs span major MCP-integrated platforms, including MCP Inspector, LiteLLM, Cursor IDE, LibreChat, and Windsurf. A systemic architectural flaw, first identified by OX Security and corroborated by the Cloud Security Alliance, sits inside Anthropic's official MCP SDKs: the STDIO transport passes parameters directly to the host operating system's shell without sanitizing them, which allows arbitrary code execution. That flaw alone exposes an estimated 200,000 instances to vulnerability across the dependent supply chain. Anthropic has confirmed the behavior is intentional and left remediation to the developers building on top of the protocol. Invariant Labs separately showed that a single malicious MCP server can weaponize adjacent, trusted servers through cross-server tool shadowing, so compromise does not stay contained to the server an attacker first targeted.

The temporal dimension compounds all of this. A document picked up during a routine web search can quietly corrupt an agent's long-term memory, and that corruption can influence behavior in a completely unrelated session weeks later, with no anomaly visible at the moment of injection or at the moment of exploitation. MCP-38, a protocol-specific threat taxonomy published in March 2026, catalogs tool description poisoning, indirect prompt injection, parasitic tool chaining, and dynamic trust violations as threat classes that existing frameworks, built for traditional software or for generic LLM deployments, simply don't cover. Each of these mechanics produces real damage. None of them produces an HTTP error.

Behavioral Baselining as a Detection Primitive

Static controls, pattern matching, schema validation, rate limiting, catch the attacks someone has already seen and described. They match signatures, and a signature only exists after an attack has been documented once. Novel attacks and evasion techniques pass straight through, because there is nothing for the rule to match against.

Behavioral anomaly detection works differently. It tracks statistical properties of agent behavior over time, not the content of any single request, and flags deviations from an established pattern. The working premise is simple: what an agent does consistently defines what it should keep doing. When that pattern breaks, something changed, whether that's a compromised tool, a coerced instruction, or a misconfigured update.

Pirch et al. work out of TU Berlin and the Max Planck Institute, and they frame this through an operating-system lens. AI agents face the same structural problems operating systems have always faced: isolating resources, separating privileges, mediating communication between components. They place the agent runtime in the role of the kernel, where it arbitrates access and enforces policy, and behavioral monitoring becomes the agent equivalent of kernel-level syscall auditing. A baseline built on that model has to capture several dimensions together: tool-call frequency, call sequencing, error rate, output token volume, and network egress. Watching any one of these in isolation misses the signal, because the signal lives in the combination and in how that combination shifts over time.

Consider an agent that always calls a version-control log command before it makes any change, and one day it stops. Or an agent that begins calling a destructive tool it has never touched before. Neither event trips an HTTP error. Both are exactly the kind of deviation a behavioral baseline exists to catch.

Building a tool-call baseline that can fire a meaningful alert

A useful baseline is a learned model, not a fixed threshold you set once and leave alone, of what this specific agent, in this specific role, calling these specific tools in this specific sequence, looks like when nothing is wrong.

An unexpected sequence of calls reveals a structured deviation that a single call cannot. A recurrent neural network, typically an LSTM, trained on sequences of agent behavioral events, tool calls, input-output pairs, temporal spacing, learns to predict what should happen next in that sequence. The prediction error becomes the anomaly signal. One unexpected call might be noise. An unexpected sequence of calls is a structured deviation, and structure is what separates a real signal from a coincidence.

Baselines also need to account for intent, beyond pattern alone. Without conditioning the baseline on the agent's current goal, a system that treats all activity as one undifferentiated behavioral envelope will throw a false positive every time the agent is handed a new task, which makes the baseline useless the first week someone changes its job. If a baseline knows the agent's assigned goal, it can separate adaptation from compromise.

Deployment events have to feed the baseline directly. A model update changes inference latency. A prompt revision changes the order in which tools get called. If the baseline doesn't know a new model version went live, it will read ordinary post-update behavior as an attack and fire on something that was never a threat. Chu et al.'s attack-surface grid makes the stakes of getting this right concrete: seven of twenty-eight cells in that grid have zero defense coverage today, and three of those seven already contain documented attacks. Several of the uncovered cells map directly to cross-session and sub-session-stack threat classes that per-call monitoring structurally cannot reach. Sequence-level baselining, tied to deployment events and conditioned on goals, is the mechanism that closes them.

Alert fatigue and chain-based emission

The strongest objection to all of this is a practical one, and anyone who has sat through a monitoring system's 3 a.m. alert queue will recognize it immediately: legitimate agent activity throws off dozens of individually indistinguishable signals per task: tool invocations, network egress, file reads, API calls. Each one looks exactly like the agent doing its job correctly, because in isolation, it is.

Tuning thresholds tighter doesn't fix this. Every CI/CD cycle that ships a new model version, adds a tool integration, or updates a prompt template shifts the agent's behavioral envelope, and nothing in its declarative configuration shows this to a human. A threshold tuned correctly yesterday is wrong tomorrow, and tightening it further just shifts which legitimate actions get flagged. Suppressing signals to cut the noise is worse. A prompt-injected tool call has exactly the same shape as a normal tool call until the surrounding chain gives it context, so removing the individual event to reduce volume removes the only evidence that would have revealed the attack.

The fix is structural: change the unit of detection from the individual tool call to the causal chain of tool calls. Per-event telemetry across endpoint, network, and agent layers generates a volume of true positives that no downstream triage process can compress into something a human can act on. Collapsing that volume at the chain level preserves the attack evidence while cutting the noise that drowns it. Chu et al. note that no benchmark surveyed in their work evaluates cross-session or sub-session-stack threats at all, so progress on exactly the threat classes that require chain-level detection is currently unmeasurable by any existing standard. Building chain-based detection now, on the architecture the threat demands, is the argument, rather than waiting for a benchmark to catch up to an architecture it happens to score.

Identity and authorization as preconditions for baseline accuracy

A behavioral baseline only works if the thing being baselined has a coherent identity. Without one, anomalous behavior from a single compromised agent instance can hide behind normal behavior from a different instance that shares the same credentials, so the baseline never separates the two.

AI agents need an identity class of their own, distinct from a human user and from a traditional service account, because an agent's authentication carries a delegation chain: this agent acts on behalf of this user, with this scope, for this purpose. Neither a human identity model nor a service-account model captures that chain cleanly, and forcing an agent into either one produces a baseline that can't tell the agent's behavior apart from the human user's. Several operational risks make this worse in practice. If an agent is issued broad, read-everything or write-everywhere tokens, you cannot define a meaningful normal access pattern in the first place, because under that scope, everything the agent touches is technically permitted. Credentials stolen from agent runtime memory during a prompt injection or a supply-chain compromise get reused under the original agent's identity, so the baseline may see the reuse and read it as the same agent behaving normally. And prompt-injection coercion is the hardest case of all: an attacker injects instructions that push the agent to use its own legitimate scopes for a malicious purpose, so the credential checks out, the scope checks out, and only the intent behind the action is adversarial. Behavioral baselining catches that last case because nothing about the access control itself was violated.

Least-privilege scoping is a precondition for a baseline that can fire a meaningful alert at all, because a token scoped to everything makes every action look like something the agent was supposed to do.

What audit trails must capture for governance evidence

A baseline alert tells a security team that something changed. It doesn't tell them who authorized the action that triggered it, what context the agent had available at the moment it made its decision, what it actually decided to do, or whether that decision lined up with policy. Without those four answers, an alert is a signal that something is worth investigating, not evidence anyone can act on or defend later.

Agentic deployments introduce governance gaps that make it harder than it sounds to capture those four answers. Risk can accumulate across sessions in ways that stay invisible until it's substantial. An agent can be authorized for an action in general but not for the specific context it was applied in. An agent's scope of action can change at runtime, mid-session, and its static configuration never shows that change. And logging what happened after the fact proves nothing if there's no record of what policy governed the decision at the moment the agent actually made it.

Tamper-proof records make an agent's behavior, which is probabilistic and context-dependent, harder to dispute after an incident than a mutable log would allow. A cryptographically signed record of what the agent executed, and under which policy, is much harder to dispute. Full OpenTelemetry tracing closes the remaining gap between the behavioral signal and the causal chain behind it: when a baseline alert fires, the trace has to show the full sequence of tool calls, the inputs the agent received at each step, and the outputs it produced, not merely the fact that a call was made.

The timeline for building this capability is no longer open-ended. The EU AI Act's high-risk obligations apply from December 2, 2027 for stand-alone systems, following the Digital Omnibus, agreed in May 2026 and in force as of July 27, 2026. Enterprise agentic deployments operating in Annex III use areas, recruitment, credit scoring, essential services, qualify as high-risk under the current definition, while deployments in lower-stakes contexts do not. Enterprises now have a fixed horizon to build audit capability against, not an open-ended best practice to get to eventually.

A Platform Implementing All of This, in Practice

A platform built to this standard starts with identity: every agent carries its own delegation chain, scoped to the narrowest set of permissions its task requires, so that the baseline built on top of it has a meaningful definition of normal to work from. Detection happens at the level of the causal chain, not the individual event, so the system surfaces the handful of chains that actually warrant review instead of the hundreds of individually unremarkable calls that make up a normal day. And every alert that does fire carries a full OpenTelemetry trace and a tamper-proof record behind it: who authorized the action, what the agent saw, what it decided, and whether that decision matched policy, ready for a security team to act on and for a regulator to review.

Each one of these pieces, identity, sequence-level baselining, chain-based emission, auditable traces, addresses a gap that latency dashboards and error-rate monitors were never built to see. Put together, they close those gaps. If you leave them apart, each one leaves the kind of blind spot the attacks above are already built to exploit.

Sources

  1. A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
  2. Toward Securing AI Agents Like Operating Systems
  3. MCP Security Crisis: Systemic Design Flaws in AI Agent Infrastructure
  4. MCP-38: A Comprehensive Threat Taxonomy for Model Context Protocol Systems (v1.0)
  5. Model Context Protocol (MCP): Security Design ...
  6. Securing AI agents: When AI tools move from reading to acting

More in Agentic Incident Response