Agentic AI Traces: Defend What You Can Replay

Agentic AI traces newspaper-style header with the thesis "You cannot defend a decision you cannot replay. You can trace every one." and three Thinkers360 Certified Expert badges, by Dr. Harish Kotadia, Ph.D.

What are agentic AI traces?

Agentic AI traces are the linked record of one agent run. They tie the first prompt, every model call, every tool call, and every permission decision to one ID, so the run can be replayed end to end. A log tells you something happened. A trace tells you what led to it. So when an agent’s decision is questioned, the trace is the only defense that does not rest on memory.

New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.

You cannot defend a decision you cannot replay. You can trace every one.

Why do agents fail without traces?

Because nobody can see why. Anthropic’s multi-agent research post puts it plainly: “Agents make dynamic decisions and are non-deterministic between runs, even with identical prompts.” Users reported agents “not finding obvious information,” and the team “couldn’t see why.” Then they added tracing, and in their words, “Adding full production tracing let us diagnose why agents failed and fix issues systematically.”

I read that line as the whole case. A loop that runs forty steps and picks a new route each time cannot be debugged from the final answer. Still, most teams I meet log the answer and nothing else. When something breaks, they re-run the prompt, get a new route, and call it flaky.

There is a second reason, and it matters more in my world. A reviewer asks why the agent did X. Without agentic AI traces, the honest answer is “we think it was because of Y.” That is not a defense. It is a guess with a timestamp.

What does a trace have to link?

Four things, under one ID.

  • The prompt that started the run.
  • Each model call it triggered, with the tokens and the model used.
  • Each tool call, with what it did and whether it worked.
  • And each permission decision, with who or what made it.

If any of the four is missing, the replay has a gap. The gap is always where the question lands.

The ID is the part teams skip. Four separate logs, each correct, are not a trace. They are four piles. So the first design choice is one ID on every event, from the first prompt to the last tool result. Without it, the rebuild is by hand, and a rebuild by hand is where auditors stop trusting you.

How do the vendors export agentic AI traces?

Claude Code exports through OpenTelemetry, and the monitoring docs are specific about what comes out. Every event carries a prompt.id. The docs define it as a “UUID v4 identifier linking all events produced while processing a single user prompt.” That is the one ID I asked for above, and the vendor supplies it. The tool_decision event records the decision and its source, which can be config, hook, or one of the user options. So the trace shows not just that a tool ran, but who let it.

Spans are newer. The docs say the tracing beta exports “spans that link each user prompt to the API requests and tool executions it triggers, so you can view a full request as a single trace in your tracing backend.” It sits behind a flag as I write. But the events and the prompt.id do not, and they carry most of the evidence already.

Two details make this a governance control, not a developer perk. First, content is off by default: prompt text is <REDACTED> unless an admin turns it on. Second, where it goes can be locked. Set the exporter in managed settings and, as the docs put it, “Claude Code removes conflicting developer-set variables at startup.” A repository’s own settings file cannot turn telemetry on or point it elsewhere. So the trace goes where the company says, not where the developer says.

On the standards side, the OpenTelemetry GenAI conventions define invoke_agent and execute_tool spans, with gen_ai.agent.name, gen_ai.tool.call.id and gen_ai.conversation.id as fields. They are marked Development, so the names may still change. Content capture is Opt-In there too. I use the names anyway, because a moving standard beats a private one.

Here is how the pieces map.

Trace element Where the harness supplies it Advisory or enforcing
One run ID prompt.id on every event; session.id Enforcing (runtime writes it)
Model calls api_request, api_error events; spans in beta Enforcing
Tool calls tool_result event with duration_ms, success, error Enforcing
Permission decisions tool_decision with decision and source Enforcing
Destination OTEL_EXPORTER_OTLP_* in managed settings Enforcing once managed; advisory if left to the developer
Content OTEL_LOG_* gates, redacted by default Policy choice, not a default

More on the Agentic AI Evidence Layer



What do I turn on first?

The ID and the endpoint, before anything else. If prompt.id reaches a collector the company owns, the rest is a query. So my first deliverable is one managed settings entry, with the exporter and endpoint, pushed to every machine that runs an agent. The developer can add local logging. They cannot turn the company trace off.

Second, permission decisions. I want tool_decision with its source in the collector from day one, because that is the row a reviewer reads. Third, tool results with success and error. Fourth, content, but only once the privacy question is answered in writing. Anthropic’s team watched “decision patterns and interaction structures” without reading the chats. That is where I start. Fifth, a replay drill. Pick one run from last week, rebuild the route from the collector alone, and time it. If it takes an afternoon, the trace is not done.

Where does this sit in the six layers?

Agentic AI traces live in the evidence layer of my six-layer architecture, next to the eval suite and the audit trail. But they depend on the enforcement layer below. The managed setting is what makes the endpoint firm. Then the evidence layer turns that stream into something a person can replay.

On the Five-Stage Roadmap, traces are the line between Governed and Assured. Governed teams have the collector and the ID. Assured teams have run the replay drill and can show a reviewer a route in minutes. Piloted teams have console logs on a laptop. Fine for a pilot, useless for a finding.

What transfers to regulated loan origination?

The whole thing, and the reviewer already exists. In my work in regulated loan origination, every credit decision has to be rebuildable: what data came in, what rule fired, who overrode it. That is a trace by another name. So agentic AI traces are not a new ask there. They are the existing ask, applied to a system that picks its own route.

The part that changes is the level of detail. A rules engine has one path. An agent has forty steps, and a handful of them are tool calls that read or write a file. The trace has to capture each step, or the rebuild has holes. I keep content redacted and keep the structure complete. A reviewer needs to see that the agent read the pay stub and then called the underwriter. They do not need the pay stub itself.

Roadmap diagnostic: pick one agent run from last week. Rebuild its full route from your collector alone, without asking the developer who ran it. If you cannot, you cannot defend it.

The bottom line

Instructions in, results out was IT. Intent in, outcomes out is agentic AI. But an outcome you cannot replay is one you cannot defend. A governance program that rests on “we think” is not a program. So carry one identifier through the run, lock the destination in managed settings, and drill the replay before a reviewer asks for it.

My books go deeper on both sides of this. Intent In, Outcomes Out covers the architecture. Earned Autonomy covers how the evidence buys the next tier.

Book covers of Intent In, Outcomes Out and Earned Autonomy by Dr. Harish Kotadia, Ph.D., two field guides to agentic AI architecture and governance.

Could you replay last week’s agent run from your collector alone, or would you need to ask the developer?

Go deeper

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.

Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

 


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading