What is agentic AI harness drift?
In August I wrote that the harness is not adjacent to governance, it is governance implemented in code. Agentic AI harness drift is what happens when that code changes and nobody records it. The prompts, hooks, tool schemas, skills and permission rules around a model move every week, while the model version stays pinned and the change log stays empty.
New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.
You cannot claim change control because you pinned a model version. You can claim it when the harness itself is an inventoried, versioned, eval-gated artifact.
That August post made one argument. You cannot govern a frontier model, but you can govern the harness you built around it, completely and with evidence. This post takes the next step. If the harness is the thing you govern, then it needs the discipline a bank gives a model. An inventory, a version, and a test before each change goes live.
How much can the harness move an outcome?
More than most model upgrades do. In February, an agent-framework vendor held its model fixed and changed only the harness. Its agent went from 52.8 percent to 66.5 percent on Terminal-Bench 2.0, from the top 30 to the top five. The changes were a build-verify loop in the system prompt, a pre-completion checklist, a loop detector and a context map at startup. Not one line of model weights changed.
The reverse is also true. Harness pieces built for one model become dead weight on the next. An engineer on Anthropic’s platform team described context resets his team added for one Claude model that a later model no longer needed. His line is the one I keep coming back to: assumptions about what the model cannot do need to be re-tested with each step change in its capability.
So a harness change can add 14 points and an unchanged harness can quietly cost you points. Either way, if that change is not versioned, you cannot say which harness produced which outcome. That is agentic AI harness drift, and an examiner will find it before you do.
What does a versioned harness look like?
A global ride-hailing platform is the cleanest public example I have seen. In an August 27 engineering post, one of its distinguished engineers described holding one model fixed from February to July so that every gain could be traced to a harness change. The team tuned compaction thresholds, reasoning defaults, cache settings, tool search and the default model for subagents. Cost per session fell 52 percent from its June peak. Cost per thousand model requests fell almost 34 percent.
The sentence that matters for governance is a quiet one. The post said the subagent default setting proved to be the most impactful lever. A default setting. Not a model, not a prompt, a configuration value in the harness. Because the model was frozen, the team could prove that claim with a chart instead of a hunch.
That is what I mean by versioning the harness like a model. Each configuration value had a before, an after and a measured effect. In my work in regulated loan origination, that is exactly the evidence a model validation team asks for. Most agent teams cannot produce it, because the value changed in a config file on a Tuesday and nobody wrote it down.
More on the agentic AI harness
- Govern the Harness: The August post this one builds on, covering who owns the harness across the Five-Stage Roadmap.
- Credential Scoping: Why an agent gets one scoped key, never yours.
- Definition of Done: The gate that decides when an agent’s work is finished and who signs it.
- Change Management: How a regulated firm moves agents through controlled change without losing the audit trail.
What does unversioned drift look like?
It looks like permission rules written by whoever was annoyed last. Anthropic reported in its August 7 post on auto mode that 49.5 percent of active Claude Code CLI users had hand-written a Bash allow-rule by June. That share was growing about five percentage points every five weeks. Each rule is a harness change. Each one widened what the agent could run without asking, and each one lived on one developer’s machine.
That is agentic AI harness drift at the desk level, policy moving up from the user instead of down from the organization. Nobody approved it, nobody versioned it and nobody can list it. I have seen the enterprise version of this pattern more than once. A lending team ships an agent with a tight tool list in month one. By month four, three engineers have added tools, loosened a hook and raised a token budget to make a demo work. The model version in the risk register has not changed. The system has.
The contrast with the ride-hailing platform is the whole post. Same kind of harness changes, same kind of levers. One company froze the model and logged every lever. The other pattern lets the levers move in the dark.
What goes in the harness register?
Everything that changes agent behavior without changing the model. The system prompt and its version. Every hook, with its trigger and its author. Every tool schema and every skill, pinned to a release. Permission and allow-rules, including the ones developers add locally. Compaction thresholds, reasoning defaults, token budgets and subagent defaults, the levers the ride-hailing team tuned. And the eval result that each version produced before it went live.
That last line is the gate. Two engineers at a major cloud vendor wrote on September 9 that an eval suite is not there to celebrate a 2 percent gain. It is there to give you confidence that a prompt tweak, tool schema change or model upgrade did not make the agent worse. They also warn against gating on a single noisy run, so batch the evals and track the pass rate over time. Either way, no harness version ships without a score next to it.
I would add one more column, borrowed from that Anthropic engineer: what can I stop doing? Each new model is a reason to re-test the scaffolding and retire the pieces that have gone stale. A harness register that only grows is its own kind of agentic AI harness drift.
Where does this leave the risk committee?
Instructions in, results out was IT. Intent in, outcomes out is agentic AI. In that world the model is the least changeable part of the system and the harness is the most changeable. Governing the pinned part and ignoring the moving part is not change control. It is a model number on a form.
Version the harness like a model. Inventory it, pin it, eval it, and retire what has gone stale. The August post said the harness is governance implemented in code. This one says the obvious next thing: code has versions, and so should your governance.
Does your model inventory list the harness, or only the model it wraps?
Was this post useful? One click tells me what to write next.
Go deeper
- Agentic AI Architect: control design for enterprise agents.
- Agentic AI Case Studies: deployment evidence, one teardown at a time.
- Agentic AI P&L: cost, payback, and risk in a CFO’s voice.
- Agentic AI SDLC: the lifecycle that builds the agents.
- Subscribe to the Blog: I publish a new post every weekday on one agentic AI governance question. Enter your email and each post arrives in your inbox the moment it goes live.
© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026
Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.
Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.
Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

