What is an agentic AI incident record?
An agentic AI incident record is the account of one agent run that went wrong, written while the run is still going. It captures the failing step, the state at that moment, and the steps that led there. It also captures the decision to stop or resume. A postmortem written the next day rests on memory. So the agentic AI incident record is the one that does not.
New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.
You cannot explain a loop after the fact. You can record it during.
Why can I not explain a loop after the fact?
Because the loop does not keep its own history. An agent runs forty steps. It compacts its context near step twenty. Next time it picks a new route. By the time someone asks what happened, the evidence is a summary, or gone, or was never written at all.
Anthropic showed what this costs. In its postmortem of three production issues, the team wrote that its evaluations “simply didn’t capture the degradation users were reporting, in part because Claude often recovers well from isolated mistakes.” Then the harder line: “we lacked a clear way to connect these to each of our recent changes.” Privacy controls kept engineers away from the affected conversations. So the rebuild took weeks.
I read that as the general case, not an Anthropic case. A team with that much tooling could not explain a fault after the fact. A bank with console logs on a laptop has no chance. The harness has to write the record during the run, because the run is the only witness.
What does an incident record have to capture?
Five things, and the first one is a trigger. The record starts when a tool call fails, an API call errors, or a turn ends badly. Not when a human notices.
Then the state at that moment: which prompt was live, which model, what the context held. Then the chain since the last good step, because errors in agents compound. Anthropic’s multi-agent research post puts it in four words: “Agents are stateful and errors compound.” A failure at step thirty usually started at step twelve.
Fourth, the call: stop, retry, or resume. The same post says the team “built systems that can resume from where the agent was when the errors occurred,” because restarting from zero “is expensive and frustrating.” Whether the run resumed, and from where, goes in the record. Fifth, who was told and when. If nobody got paged, it is not an incident yet, and a reviewer will ask why.
How do the vendors supply it?
Claude Code fires a hook at each of those five moments, and the hooks reference names them. PostToolUseFailure runs “after a tool call fails.” StopFailure runs “when the turn ends due to an API error.” SessionEnd runs when a session terminates. It carries a reason field, so the record can say whether the operator cleared it, resumed it, or logged out. Every hook receives the session_id, the prompt_id, and a transcript_path.
PreCompact is the one I care about most. It fires “before context compaction,” the last moment the full loop history exists in one place. A hook there can copy the transcript to storage the company owns. After that, the agent’s memory of the incident is a summary it wrote about itself.
One caveat from the same docs, and it matters. The transcript file “is written asynchronously and may lag the in-memory conversation, so it may not yet include the current turn’s most recent messages when a hook fires.” So the hook writes the event fields it is handed first, then fetches the transcript a beat later. The monitoring docs add the api_error event, with status code and retry count. It lands in the same collector as the traces.
Here is how the five parts map to what the harness gives you.
| Incident element | Where the harness supplies it | Advisory or enforcing |
|---|---|---|
| Trigger | PostToolUseFailure, StopFailure hooks; api_error event |
Enforcing (runtime fires it) |
| State at failure | Hook input: session_id, prompt_id, permission_mode, transcript_path |
Enforcing |
| Chain since last good step | Transcript snapshot at PreCompact; trace events by prompt.id |
Enforcing once the hook is managed |
| Stop or resume decision | SessionEnd with reason; Stop hook exit code |
Enforcing |
| Who was told | Notification hook to a pager or channel |
Advisory until the hook lives in managed settings |
More on the Agentic AI Evidence Layer
- Traces: The normal-run record that one ID has to link, which the incident record starts from.
- The Audit Trail: Log the handoff, not the answer, so the chain before the fault exists.
- Context Compaction: Why compaction is a deletion with a narrative attached, and why the snapshot comes first.
- The Circuit Breaker: A rehearsal, not a button, and the record shows whether it tripped.
- Ten Controls in One Place: The full governance controls roundup, ten questions and ten answers.
What do I record first?
The trigger and the snapshot, before anything else. My first deliverable is two hooks in managed settings. One on PostToolUseFailure writes the event fields to the company collector. One on PreCompact copies the transcript to the same place. Both live where the developer cannot edit them. That is the whole gap between an agentic AI incident record and a developer’s notes.
Second, StopFailure and SessionEnd, so the record has an ending. Third, the Notification hook to a channel a human reads. Fourth, a resume rule in writing: when the run may pick up from the failed step and when it must stop. I keep content redacted here too. A reviewer needs to know that step twelve read a file and step thirty failed. They do not need the file.
Last, a drill. Break one run on purpose, then hand the collector output to someone who was not in the room. If they can tell you what failed, what came before, and what happened next, the record works. If they call the developer, it does not.
Where does this sit in the six layers?
The agentic AI incident record lives in the evidence layer of my six-layer architecture, beside traces and the audit trail. Traces record the run. The audit trail records the handoffs. The incident record is what those two become when something breaks, plus the decision that followed. It leans on the enforcement layer to keep the hooks managed, and on the runtime layer to fire the snapshot before compaction.
On the Five-Stage Roadmap, this is a Governed control. Piloted teams hear it from a user. Governed teams hear it from the hook, with the snapshot already taken. Assured teams have run the drill and timed it.
What transfers to regulated loan origination?
All of it, and the form already exists. In my work in regulated loan origination, a decisioning fault gets an incident report. What failed, which applications it touched, what got rolled back, who signed off. Nobody writes that from memory a week later. The agentic AI incident record is the same form. The harness fills it in instead of a person, at the moment of the fault instead of after.
What changes is the number of steps in between. A rules engine fails at one rule. An agent fails at step thirty because of step twelve. So the chain matters more here than anywhere else I have worked. The snapshot is what keeps it.
Roadmap diagnostic: break one agent run on purpose this week. Hand the collector output to someone who was not there. Can they explain it without calling the developer?
The bottom line
Instructions in, results out was IT. Intent in, outcomes out is agentic AI. But an outcome that went wrong needs a record made while it was going wrong, not a story told later. A postmortem from memory is a guess with a signature. So wire the hooks, snapshot before compaction, and keep the record where the developer cannot reach it.
My books go deeper on both sides of this. Intent In, Outcomes Out covers the architecture. Earned Autonomy covers how the evidence buys the next tier.
When your last agent run failed, who wrote the incident record, and when?
Go deeper
- Agentic AI Architect: control design for enterprise agents.
- Agentic AI Case Studies: deployment evidence, one teardown at a time.
- Agentic AI P&L: cost, payback, and risk in a CFO’s voice.
- Agentic AI SDLC: the lifecycle that builds the agents.
- Subscribe to the Blog: I publish a new post every weekday on one agentic AI governance question. Enter your email and each post arrives in your inbox the moment it goes live.
© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026
Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.
Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.
Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.


