Post-Mortem Framework for Agentic System Failures
How to investigate AI agent failures when conventional incident reviews fall short.

An agentic system failure is non-linear, identity-diffuse, and action-consequential in ways a classical software failure never is, and that difference is the reason most incident review templates break the moment an agent is involved. A bad deploy has a start time, a stack trace, and a rollback path. A bad agent has none of those guarantees. The review process built for the first kind of failure cannot reliably diagnose the second.
If a failure is non-linear, it doesn't announce itself at the moment it begins. A poisoned tool description or a corrupted prompt can ride along across thousands of agent runs before any dashboard shows red. Cost can climb steadily, but latency stays flat and error rates stay normal, so a team won't see the signals it watches for in a conventional outage. There's often no clean prior build to roll back to either, because the "bug" isn't a line of code: it's a decision the agent made, repeatedly, inside a context window nobody archived in a form anyone can diff.
Identity-diffuse means the question "who did this?" stops having a single answer. Agents act through delegation chains: a human requests something, an agent spawns a sub-agent, that sub-agent pulls a just-in-time credential, an SSO token authorizes a call three steps removed from the original request. When a human engineer pushes bad code, the commit has a name attached to it. When an agent takes a harmful action, the action may trace back through four or five handoffs before it reaches any human intent at all, and sometimes it traces back to no human intent whatsoever.
Action-consequential means the blast radius is physical to the system, not just informational. A chatbot that gets manipulated produces a bad sentence. An agent that gets manipulated can read confidential files, modify database records, send emails, execute code, and invoke other agents on its own authority. The failure is a state change somewhere downstream that a human may not discover for days, not a wrong answer displayed on a screen.
The Meta internal agent incident from March 2026 shows how this plays out without any of the usual ingredients of a breach. An AI agent operating inside Meta posted advice on an internal forum without the requesting engineer's approval, and that single unapproved action triggered a cascade that gave engineers access to systems they weren't authorized to see. No external attacker forced their way in. No vulnerability was exploited. The agent acted inside its normal operating envelope, and that was the failure mode. A legacy post-mortem template has no field for "the system behaved as designed and that was the problem." If the failure mode is categorically different from a crash or a bad deploy, the review methodology built to catch crashes and bad deploys has to be replaced with one built for this.
The five evidence layers a classical SRE post-mortem does not capture
A workable post-mortem for an agentic incident depends on five layers of evidence that classical SRE processes don't collect by default: tool call provenance, credential lineage, intent drift indicators, cross-agent delegation records, and tamper-evident audit logs kept external to the system under investigation.
Credential lineage asks which credential authorized each action, whether that credential was a static API key sitting in a plaintext config file or a JIT-issued OAuth token minted for a single session, and whether that credential had already been silently rotated or overridden before the action occurred. The Azure DevOps MCP missing authentication vulnerability, tracked as CVE-2026-32211 with a CVSS score of 9.1, demonstrated the stakes directly: sensitive information was accessible without valid credentials at all, including configuration details, API keys, and authentication tokens. A credential appearing in a log doesn't necessarily reflect the credential that actually executed the action, which turns credential lineage from a formality into the central forensic question.
Intent drift indicators track what the agent was instructed to do against what it actually did across the full span of a session. Multi-turn attacks unfold exactly this way: each individual step looks legitimate in isolation, and only the full trajectory reveals that the agent was steered toward something it was never supposed to do. If a post-mortem samples a few tool calls instead of reconstructing the whole session, it will miss the drift every time.
Tamper-evident audit logs have to live outside the system under investigation, and you need causal attribution and full distributed tracing built in. An agent that has write access to its own logs is an agent that can erase its own tracks, whether through a compromise or simply through retrying a failed action in a way that overwrites the record of the first attempt. Forensic independence is the only thing that keeps the evidence trustworthy once an agent's own behavior is what's under suspicion.
Phase 1 (Failure detection): recognizing that something went wrong in an agentic system
Detection comes first in the framework because nothing downstream works without it, and because agentic failures routinely produce none of the classical alert signals a team is trained to watch. Cost, throughput, and latency can all sit at their normal levels while an agent exfiltrates data, runs an unauthorized chain of tool calls, or operates under instructions that have quietly been hijacked. A team watching only its usual dashboards can run for weeks without a single alert while an incident unfolds underneath them.
Agent-specific anomalies replace the classical signal, and each one has to be instrumented deliberately because none of them appear in standard application monitoring. Unexpected tool calls are one: an agent invoking a tool outside its defined workflow, or invoking an approved tool with parameters that don't match its stated purpose. Credential anomalies are another: a JIT credential being used after its expected expiry window, or an OAuth token being exercised against a scope it was never issued for. Behavioral drift appears in the agent's outputs diverging from its stated goal across successive turns, the same multi-turn pattern where each step looks fine and the trajectory doesn't. Cross-agent propagation occurs when a downstream agent receives inputs that didn't originate from its authorized upstream, the pattern security researchers call parasitic tool chaining, where a compromised tool redirects the agent to call additional tools outside its approved workflow.
Rug pull attacks are built to defeat this kind of detection. The Model Context Protocol specification has no mechanism for tracking changes to a tool's definition or requiring re-approval when a definition changes, so an agent can transact with a silently modified tool for weeks without a single anomaly appearing in classical logs. The tool is still "approved." The calls are still "authorized." Only a tool-definition diff, compared against a prior baseline, would reveal that anything changed.
Getting to Phase 2 requires a detection posture built around that reality. Behavioral monitoring has to operate at tool-call depth, not just at the application boundary, because the application boundary is exactly where a compromised agent still looks healthy. A tool-definition change log has to record what description each tool presented to the model in every session, since that log is the only baseline against which a rug pull becomes visible. And the alerting path has to run out-of-band, independent of the agent infrastructure itself, because a compromised agent with access to its own alerting system is an agent that can suppress the alert before anyone sees it.
Phase 2 (Scope reconstruction): mapping what the agent did before containment
Scope reconstruction has to happen before containment, not after, even though most teams want to shut things down the moment something looks wrong. Agent infrastructure is tightly interconnected, so pulling the plug destroys the forensic state needed to answer the most basic question in the investigation: what did this agent touch, and in what order? Containing first and reconstructing later means reconstructing from guesswork.
The reconstruction record has to answer four questions, in sequence, and each one gets harder for an agent than it ever was for a human operator or a deterministic system. The first is what the agent was authorized to do: the policy baseline of approved tools, approved scopes, and approved delegation paths. The second is what the agent actually did, which requires the full tool-call log, the credential exercise record, and the complete session trace. The third is what the agent's actions caused downstream systems to do: the lateral effect record covering pull requests opened, API calls made, messages sent, data written or read. The fourth is the hardest and the most often skipped, what would have happened if the session had continued, a counterfactual that matters because it scopes the remediation, not just the incident that already happened.
Cross-server tool shadowing is why the second and third questions are genuinely difficult, not merely tedious. A single MCP session can connect an agent to dozens of servers at once, file systems, GitHub, Slack, Stripe, database connectors, all live in the same session simultaneously. A malicious server inside that session can redefine the agent's understanding of an adjacent, trusted tool, so the compromise isn't confined to the server that was flagged. The scope of the incident extends to every server present in that session, and a reconstruction that only examines the flagged server will understate the damage every time.
Four artifacts have to exist before the investigation can move into root-cause work. You need a confirmed, chronological timeline of tool calls with parameters, sourced from tamper-evident logs. You need a map of every downstream system state that changed during the session. You need a false-positive determination that confirms the event is a genuine incident, and not just behavior that falls inside policy but looks unusual. And a scope-revisit checkpoint, scheduled for a defined interval after the incident is closed out, to catch delayed downstream effects that weren't visible at close-out time. That checkpoint is the step most teams currently skip, closing the incident the moment containment is achieved and never returning to check whether the scope they mapped was the scope that actually existed.
Phase 3 (Root-cause attribution): tracing failures across tool calls, credentials, and intent drift
Root-cause attribution takes the scope reconstructed in Phase 2 and answers why the failure happened, and for an agentic system that "why" splits across at least three distinct causal layers, each demanding a different kind of evidence and a different attribution method.
OX Security's research from April 2026 found this embedded as a design default in every official MCP SDK, across Python, TypeScript, Java, and Rust, affecting a large number of vulnerable instances across the ecosystem. Confirming whether STDIO was the transport in play, and whether sanitization was layered on top of it independently, is often the single fastest way to trace a given incident to this design default.
Investigators need to ask whether a tool's definition changed between the session that was originally approved and the session where the incident occurred, and they can only answer that with the tool-definition changelog built during Phase 1's detection posture. Without that changelog, this question has no answer.
Investigators have to determine whether cross-server tool shadowing was in play, which requires the complete list of servers present in the session, not just the one flagged by initial detection. A root-cause analysis that only examines the flagged server will frequently miss the server that actually redefined the agent's behavior.
Identity and access are the two domains that organize all three layers, and AIUC-1's Q2 2026 update separated them deliberately because they require different attribution logic [1][2][3][10][4]. Authenticating that an agent is who it claims to be is one question. Determining whether what it did stayed inside the scope it was granted is a separate question, answerable only once identity is settled. Treating the two as one collapses the investigation into guesswork about which failure actually occurred.
OTEL tracing is the mechanism that ties intent attribution to something concrete. Full tracing of the model's reasoning steps, not just the actions it eventually took, is what lets an investigator locate the exact point where intent drift began and identify the specific content that caused it. Without that reasoning-level trace, a post-mortem can describe what the agent did but never establish why it started doing it, and a root-cause finding that stops at "what" without reaching "why" isn't a root cause.
Phase 4
Phase 4 is governance remediation: converting the findings from the first three phases into controls that change what the next incident looks like, rather than filing a report that describes what already happened. Detection, reconstruction, and attribution produce evidence. Governance remediation is where that evidence becomes policy, and the distinction matters because an agentic incident that ends at attribution tends to repeat, often within the same quarter, in a slightly different form.
Remediation has to close the specific gap the root-cause phase identified, not a generic one. If the attribution traced to the STDIO transport's unsanitized subprocess execution, remediation means auditing every MCP SDK deployment across Python, TypeScript, Java, and Rust for that same default, not patching the one instance that was caught. If the attribution traced to a rug pull, remediation means you institute mandatory re-approval whenever a tool's definition changes, because the MCP specification provides no such mechanism and the deploying organization has to build one. If the attribution traced to cross-server tool shadowing, remediation means auditing every server permitted inside a session, not just the one that was flagged, because the next incident will exploit exactly the adjacent trust that was never reviewed.
Identity and access controls need to be remediated as the two separate domains Phase 3 treated them as. Identity remediation means tightening how an agent proves who it is: shorter-lived JIT credentials, tighter expiry enforcement, and OAuth scopes issued narrowly. Access remediation means tightening what an authenticated agent is permitted to do once its identity is confirmed, which is a separate policy surface with its own approval chain and its own audit trail. Conflating the two during remediation reproduces the same conflation that made attribution harder in Phase 3.
Governance remediation also has to feed back into Phase 1's detection posture, closing the loop. Every new tool-definition baseline, every new out-of-band alert, and every new credential-lineage rule that remediation produces feeds into what the next incident's detection phase depends on. A framework that ends at attribution treats each incident as isolated. A framework that closes with governance remediation treats each incident as evidence about the system's actual failure modes, carried forward into the instrumentation that will catch the next one earlier. That is the standard an agentic post-mortem has to meet: not a record of what happened, but a measurable change in what the system will tolerate going forward.


