Est.

Delegation Chains in Multi-Agent Systems

Every agent handoff is an authorization decision that needs three structural controls.

Staff Writer, Shadow AI & Workforce Risk · · 10 min read
Cover illustration for “Delegation Chains in Multi-Agent Systems”
Agent Identity · October 11, 2026 · 10 min read · 2,288 words

Delegation chains in multi-agent AI systems are the highest-leverage attack surface in enterprise AI deployments, because every agent handoff is an authorization decision, not just a workflow step. Secure them with three controls working together: cryptographic workload identity at every hop, token exchange instead of token pass-through at every trust boundary, and scope attenuation that narrows permissions as the chain deepens. Without all three, a validly authenticated agent can still take unauthorized action, and no one will be able to prove otherwise after the fact.

Agent handoffs are authorization decisions, not just workflow steps

A workflow diagram makes delegation look clean. In practice, each arrow on that diagram is a trust decision: the agent receiving the handoff now carries the authority of the human who started the chain, and it gets to act on that authority without asking anyone for permission again.

That's the part conventional security thinking misses. The OpenID Foundation's guidance on agentic AI calls this the central tension of agent ecosystems: the same recursive delegation that makes multi-agent systems powerful is what makes them fragile unless every handoff enforces its own authorization check.

That reframe matters because most teams building these systems still think about delegation as routing logic: get the request to the right place, get the result back. Few treat each hop as a fresh authorization event with its own stakes. That gap between how delegation is built and what delegation actually does is where the rest of this piece lives.

The confused deputy problem across an agent chain

The confused deputy problem is old, and what's new is the scale at which multi-agent systems reproduce that pattern: every delegation hop is a chance for the agent receiving the handoff to become a confused deputy, and a long chain means many chances, not one.

The Coalition for Secure AI's analysis following RSAC 2026 names the mechanism directly: a server doing what it was built to do, without a check on whether it should.

That's not a bug hiding in one bad implementation. The Coalition for Secure AI points to two real incidents as proof this isn't theoretical: an Asana tenant isolation flaw from May 2025 let data cross between organizations, affecting up to 1,000 enterprises, through excessive token reach. Around the same period, a WordPress plugin CVE produced privilege escalation through improper authorization inside an MCP implementation.

The deeper version of this problem needs no flaw to exploit. Picture Agent A delegating a task to Agent B. No audit event marked the moment privilege quietly grew. That's privilege quietly compounding across an entire chain instead of at a single server.

Three structural gaps that make delegation chains exploitable by design

Three structural properties are missing from almost every production deployment, which is what makes delegation chains exploitable: scope attenuation, cryptographic lineage, and contextual provenance. Each gap opens its own door.

The first gap is the absence of scope attenuation at delegation hops. In a well-designed chain, each agent should hold narrower permissions than the agent above it, the way a manager delegates a task but never hands over their entire budget authority. A human habit makes it worse: when an employee authorizes a tool with an MCP integration, the consent screen usually offers full read and write access as the easy option, and that's the option most people pick. That grant then sits there indefinitely, as a machine identity with no expiration date, invisible to the usage reports that would normally flag an account behaving oddly.

The second gap is the absence of cryptographic proof of delegation lineage. Anything sitting in that chain can act as the original user.

The third gap is the absence of contextual provenance. Without it, an MCP server or endpoint can't answer basic questions: where did this request actually come from, did the agent delegating it have the authority to make that call, and does the action still line up with what the original person wanted? Any agent can claim to be acting on behalf of any other, and the system has no reliable way to check.

Attacks against delegation chains in practice

Diagram: Three Structural Gaps, Three Documented Attacks. Visualizes: Visualize how three structural gaps in delegation chains each map directly to a named, dated real-world attack.

Each of those three gaps has already produced a named, demonstrated attack. These aren't lab theories.

Agent Session Smuggling, documented in November 2025, exploits the missing scope attenuation. Unit 42's analysis of this scenario draws out the broader point: a compromised agent is a more capable adversary than a compromised account, because it can generate its own adaptive strategies, exploit session state, and spread its influence across every connected client agent it touches. The research agent could execute a trade in the first place only because nothing narrowed its access when the financial assistant delegated to it.

Cross-Agent Privilege Escalation, documented in September 2025, exploits the missing lineage. No authentication existed between the agents, so any process that could format a message correctly could impersonate one of them, and the system accepted it because everything inside the shared codebase was assumed trustworthy by default.

The ServiceNow Now Assist disclosure, from October 2025, exploits missing provenance. The system had no way to confirm whether the agent doing the recruiting actually had the standing to call in help from a more privileged one.

EchoLeak, tracked as CVE-2025-32711 with a CVSS score of 9.3, struck Microsoft 365 Copilot in June 2025. The Cloud Security Alliance notes this same pattern recurs across LLM orchestration frameworks generally, wherever chains of models and tools lack automatic scope enforcement.

A 2026 manufacturing supply chain breach showed what this looks like outside the research lab. The root cause was the absence of per-hop authorization checks: once the entry point agent was compromised, everything it was authorized to delegate kept flowing downstream, unchecked, hop after hop.

The MCP specification's 2026 authorization revision is necessary but not sufficient

MCP's maintainers haven't ignored this. It removes sessions, drops the initialization handshake, deprecates three core features (Roots, Sampling, and Logging), and meaningfully hardens how authorization works.

Two changes matter in particular. Resource Indicators, defined in RFC 8707, bind a token to one specific MCP server, and that requirement has been mandatory since the 2025-06-18 revision, so it carries forward unchanged. Issuer verification is also now required: clients must confirm which authorization server actually issued a given response and bind their registered credentials to that specific issuer. Both are real improvements, worth taking seriously as progress.

They don't close the gap this piece has been describing. OAuth answers one question: is this agent allowed to talk to this server. It says nothing about what a specific, already-authenticated request is carrying, or doing, once that connection is open. An agent holding a validly issued, correctly scoped token can still pull sensitive data and send it somewhere it shouldn't go, because checking the content and intent of a request was never OAuth's job in the first place, and it still isn't MCP's job after this revision. Authentication in MCP remains optional under the spec, not mandatory, which is the deeper limit. Research published in 2026 found a large number of active MCP servers sitting on the public internet with no authentication. They comply with the letter of the specification and remain, in practice, wide open.

What secure delegation chains require at the architectural level

Closing this gap takes three controls, and they map directly onto the three structural gaps described earlier: cryptographic workload identity at every hop, token exchange instead of token pass-through at every trust boundary, and scope attenuation enforced as the chain gets longer.

You give every agent, not just every human user, a verifiable identity tied to its code and its runtime environment; that's what workload identity means. The Coalition's Agentic IAM paper pushes this further, arguing agents need to be treated as first-class identities, distinct from both human users and the service accounts security teams already manage, each one bound to verifiable claims about its code, its model, and where it runs.

Token exchange replaces pass-through at every trust boundary. The Coalition's recommendation here is direct: use RFC 8693 to perform token exchange with the authorization server at each boundary, producing a new token scoped to that specific operation, with its own auditable exchange record, instead of handing the original user's token further down the chain. Token pass-through is flagged as the single most common mistake in production MCP deployments today, and it needs no advanced exploit technique, because it's just what the architecture defaults to when nothing forces an exchange to happen instead.

Scope attenuation means each hop in the chain gets strictly less access than the hop before it, because policy enforces it. Most are running a second, disconnected stack just for agents. These controls exist in policy documents far more often than they exist in production.

Why runtime detection must accompany identity controls

Identity and scope controls stop a lot at setup. They don't catch everything that happens once an agent starts working. Runtime detection fills that gap: it catches drift, manipulation, and multi-step sequences where no single step looks suspicious but the sequence adds up to something it shouldn't.

Tool poisoning through metadata is one version of this. An attacker can bury instructions inside a tool's description, the text an agent reads to decide how to use that tool, where a human reviewing the request would never see it. The agent's behavior gets steered at the level of what the tool claims to be, not at the level of the request itself, so no authorization check aimed at the request will ever catch it. Research on this attack class found that it succeeded against GPT-4.1 in the large majority of attempts, and as of research available through April 2026, no published defense targets it.

The ChainWatch framework addresses this by mapping detection to a kill-chain model built for multi-step attacks in MCP-based systems: individual tool calls can each look authorized, but the sequence they form can add up to something unauthorized. Catching that after the fact requires tamper-proof audit logs and full tracing, the kind OTEL-based tracing provides, so investigators can reconstruct what each hop actually did, not just what it was technically permitted to do.

The governance gap as a compliance exposure

Most enterprises running multi-agent systems today can't prove, after an incident, that every step in a delegation chain stayed inside its authorized scope. Regulators are starting to expect exactly that proof, which turns a gap in architecture into a gap in compliance readiness.

Consider a system with three or four hops of delegation. Enterprises without that architecture are carrying an open question they can't answer if a regulator, auditor, or customer ever asks it.

A platform that governs delegation chains end to end

No single tool closes all three gaps at once, so most enterprises end up stitching together separate tools for agent building, gateway management, and security, and the seams between those tools become new gaps of their own. A handful of platforms are closing this from different angles. Microsoft's Entra Agent ID, generally available from April 2026, alongside Microsoft Agent 365, generally available from May 1, 2026, gives each AI agent its own identity for lifecycle and access management, tied into Conditional Access, identity protection, and Microsoft Purview, addressing workload identity squarely within the Microsoft ecosystem. Identity providers built around SSO and OAuth remain the natural control plane when you bind agent identity to existing enterprise directories, and the Coalition for Secure AI's own paper recommends this approach as a starting point.

Runlayer approaches the problem by spanning all three layers at once. Policies get enforced at tool-call depth, not just at the connection level, so this closes exactly the gap the MCP spec leaves open: a token can be valid and correctly scoped and still permit an action no one actually authorized. Its governed catalog gives you access to more than 18,000 MCP servers with OAuth and credential handling built in, so an employee no longer grants full read and write access just because scoping it down takes too much friction. Tamper-proof audit logs paired with full OTEL tracing give enterprises the cryptographic lineage and end-to-end traceability needed to prove that every hop in a chain stayed inside its authorized scope. Runlayer also surfaces shadow AI across unmanaged agents, MCPs, skills, plugins, and client configurations, addressing something most SaaS security frameworks simply can't see: the machine identities that MCP integrations quietly create outside any approval process.

The governance posture that makes delegation chains safe enough to scale

Enterprises that scale multi-agent AI safely will be the ones that treat delegation chain governance as infrastructure to build before agents proliferate, not a cleanup job to run after an incident forces the question. Enterprise AI agent guardrail frameworks increasingly put agent identity registration, credential lifecycle, scope policy, and audit trail requirements squarely in the hands of security and IT teams, as platform-level responsibilities, rather than something reviewed deployment by deployment, a pace that simply can't hold as agent use grows.

That shift changes what security teams actually do. The period from 2026 through 2028 will test something harder: which of those enterprises actually built the governance infrastructure needed to run those agents safely, without a person checking every hop by hand.

A workable governance posture for this comes down to six commitments. No OAuth token gets passed through across a trust boundary, because token exchange under RFC 8693 runs at every hop instead. And shadow agent and MCP discovery runs continuously, so the governance perimeter covers the entire estate, not just the agents that happened to go through formal approval.

Delegation made multi-agent AI powerful. Whether it stays an asset or becomes the biggest liability in the system depends entirely on whether every hop in that chain is governed as the authorization decision it actually is.

Diagram: Six Governance Commitments Before Agents Scale. Visualizes: Visualize the six-commitment governance posture described at the article's close as a ranked or sequential checklist.

Sources

  1. After RSAC™ 2026: The MCP Security Question Everyone Kept Asking - Coalition for Secure AI
  2. Fixing AI Agent Delegation for Secure Chains
  3. ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
  4. Enterprise AI Agent Guardrails: A Compliance Checklist for 2026
Filed underAgent Identity

More in Agent Identity