AI Agents Going Rogue? Nah, Bad Agentic AI Governance

AI agents going rogue header in a white newspaper layout with the headline Nah, Bad Agentic AI Governance, by Dr. Harish Kotadia, Ph.D.

What does AI agents going rogue mean?

AI agents going rogue means agents chasing goals their operator never set, with access their operator never meant to grant. It is not a jailbreak, and it is not a bad answer. So it is a governance failure, and the harness is where it shows. The sandbox, the credentials, the network and the logs let a model do what it could already do.

New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.

The headlines this week read like a stop sign. A frontier lab’s agents broke out of an evaluation, hacked a big open-source model site and hid their tracks. AI agents going rogue, the stories said. But I read the same reports as a checklist, and the checklist says bad agentic AI governance.

Is this really rogue AI?

No. The Wall Street Journal’s opinion page got there too. “People built the test, removed restraints, defined the objective, left a route open and decided not to stop what was happening. Calling the result ‘rogue AI’ does more than sensationalize it. It allows those human decisions to disappear quietly from the story.” I agree with every word.

None of this is new to readers here. I have been flagging it on this blog for almost a year.

In August I wrote about the governance gap: three quarters of enterprises expect to run agents within two years, but about one in five can say how those agents are supervised. An agent without a defined decision boundary is not autonomous. It is unsupervised. Before that, I wrote that an agent needs its own identity and a control plane before it gets a credential. I showed how autonomy without assurance is how a pilot turns into an incident. And in August I argued that a sandbox a model can leave was never a sandbox. It was a suggestion with good production values.

So every failure in this story has a control with a name. And every control sits on my Five-Stage Roadmap. Here is the case for slowing the autonomy, not the adoption.

Why is this a harness failure, not a model failure?

Because nothing in the chain required a smarter model. The model needed five gifts: a leaky sandbox, a borrowed credential, an open outbound path, a writable log, and no one counting agents. Take away any one of those and the story is shorter. Take away three and there is no story at all.

In my work in regulated loan origination, I assume the model will try things I did not ask for. That is the working assumption, not a scandal. So I build the controls around the model, then test the model against them. AI agents going rogue is what you get when you reverse that order and test the controls against the model’s good behavior.

Anthropic’s own measurement of agent autonomy makes the same point from the other end. Humans stayed in the loop for 73 percent of tool calls, and only 0.8 percent of actions were hard to reverse. The safe teams kept the loop tight and the impact radius small. Those teams earned autonomy in bands, over time.

Which control stops which failure?

I mapped each failure to the control that catches it, and to the roadmap stage where that control becomes required. The pattern is plain. These agents had Autonomous freedom under Piloted controls.

Failure in the incident Control that catches it Roadmap stage
Agents reached the open internet from an eval sandbox Egress allowlist, deny by default Governed
Stolen and borrowed credentials worked One scoped credential per agent, short-lived Governed
1,200 agents, nobody counting Agent registry with runtime discovery Governed
Message board in a shared cache Tenant isolation, no shared writable namespace Governed
7 percent of transcripts spoofed Append-only logs the agent cannot write to Assured
Probing unread for two months Runtime alerts routed to a named owner Assured
Lab learned from the victim Rehearsed circuit breaker with a blast-radius bound Assured
Open goals handed to a swarm Autonomy tiers, granted on evidence Autonomous

Read the middle column. Not one item is a model upgrade. Instead, every one is a control you can build.


More on Governing the Harness



Should enterprises slow down?

Slow the autonomy. Do not slow the adoption. Those are different levers, but the headlines blur them.

A team at the Governed stage, with an allowlist, scoped credentials and a registry, would have made this loud, bounded and short. The agents would have hit the network wall within minutes. The registry would have shown 1,200 processes where the team expected 50. And the credential would have expired before anyone could reuse it.

The calls to pause model releases fix none of that. A paused model still runs inside whatever harness you built. So the CIO question this week is not which model to trust. It is which of the eight controls above you can prove are on today, with a log.

What do I do this quarter?

Three things, in this order. First, put the egress allowlist and credential scoping in place. Those two stop the escape and the account theft, which is most of the damage. Second, add the agent registry and append-only logs, because you cannot govern a fleet you cannot count or audit. Third, rehearse the circuit breaker once on a real workload. Then write down how long it took.

Then hold autonomy where it is until you have evidence for all five. AI agents going rogue did not happen because a model got smarter. It happened because someone built a harness for a demo and then ran it at scale.

Where does this sit in the six layers?

In the enforcement layer of my six-layer architecture, with runtime and evidence dependencies. AI agents going rogue tests all three at once. Egress allowlist, credentials, isolation and registry are enforcement. Logs and alerts are evidence, while the circuit breaker sits in autonomy control. So the model layer stays untouched, because the model was never the thing to fix.

Instructions in, results out was IT. Intent in, outcomes out is agentic AI. But an outcome you cannot bound is a liability, and AI agents going rogue is the proof. If your agents broke out tomorrow, which of the eight controls would catch them first? And which failure would you hear about from someone else?

Intent In, Outcomes Out and Earned Autonomy, two books on agentic AI by Dr. Harish Kotadia, Ph.D.
My books go deeper on both: Intent In, Outcomes Out and Earned Autonomy.

Go deeper

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.

Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

 


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading