Bill Gates published a roughly 6,000-word essay this week arguing that even under the best circumstances, the transition to the AI era will be one of the most turbulent periods in human history — and that there is no plan for the social, political, and economic upheaval it will cause. His framing is binary: AI becomes either the greatest equalizer ever invented or the worst source of injustice, depending on choices made in the next few years. His prescriptions are civilizational — new domestic and global institutions, “Human Reserved” job categories, a tax on AI and robots. I read the essay agreeing with the diagnosis and thinking about a much smaller version of the same problem, one that doesn’t need a new international body to fix. Inside the enterprise, the “no plan” problem already has a name, a number, and a fix that any CIO can start today. The numbers first.
Gartner expects at least 15% of day-to-day work decisions to be made autonomously through agentic AI by 2028, up from effectively zero in 2024, and it expects agentic capability built into a third of enterprise software over the same stretch. I read that stat three or four times before it landed. Not because the number is shocking — it’s the direction I’d assumed for two years — but because of what sits next to it. Deloitte surveyed 3,235 leaders across 24 countries and found that only 21% of organizations have a mature governance model for agentic AI. Three quarters expect to be running agents within two years. About one in five can actually say, with a straight face, that they know how those agents are supervised.
That’s the gap I want to focus this post on.
Why the 21% number is the one that should worry you
Deloitte didn’t leave “mature governance” vague. Its report defines it as three things working together: clear boundaries on which decisions an agent can make on its own versus which need a human, real-time monitoring that catches anomalous agent behavior as it happens, and audit trails that capture the full chain of what an agent actually did. Strip that down and it’s a short checklist. Most firms are failing at least one leg of it.
I’d put it plainer than Deloitte does. An agent without a defined decision boundary isn’t autonomous — it’s unsupervised. Those are not the same thing, and enterprises keep treating them as synonyms.
The corroborating data isn’t kind either:
- McKinsey’s 2026 AI Trust Maturity Survey put average Responsible AI maturity at 2.3 out of what I’d read as a low ceiling — only about a third of firms hit maturity level 3 or higher on governance and agentic AI oversight specifically.
- Gartner separately projects the average Fortune 500 enterprise will run over 150,000 agents by 2028, up from fewer than 15 in 2025, while only 13% of organizations believe they have adequate governance for that scale in place.
- IBM’s 2025 Cost of a Data Breach Report found firms with high shadow-AI exposure paid an extra $670,000 per breach on top of the $4.44 million global average, and 63% of breached organizations lacked any AI governance policy at all.
Put those three together and the pattern isn’t “agentic AI is risky.” It’s “agentic AI is being scaled by people who haven’t finished writing down who’s allowed to do what.”
What breaks when governance lags deployment
I want to walk through a few documented failures, not to pile on the companies involved — every one of them is a serious operation that moved fast on a genuinely useful technology — but because each one maps to a specific, fixable governance gap.
A major North American airline learned this the hard way when its support chatbot invented a bereavement-fare policy that didn’t exist. A customer relied on it, booked accordingly, and sued. The tribunal’s ruling rejected the airline’s argument that the bot was a separate entity responsible for its own statements — it called that submission “remarkable” and noted a company is responsible for everything on its website, chatbot included. Small dollar amount. The governance signal underneath it was not small: nobody owns the output of your agent except you.
A European buy-now-pay-later fintech cut roughly 700 customer-service roles and had its assistant handling two-thirds of chats within a year, publicly framing it as the equivalent of 700 full-time agents. By May 2025 its CEO told Bloomberg the company had gone too far — cost had become too dominant a factor, quality had slipped, and the firm began rehiring humans into a hybrid model. What I find most instructive isn’t the reversal. It’s that the dashboard the company was watching — resolution rate, tickets per hour — never showed the problem. Volume metrics hid a quality decline that only showed up once someone bothered to look past the average.
A top-tier global strategy consultancy’s internal AI platform was compromised in a controlled red-team exercise; the attacking agent gained broad system access in under two hours. That’s the timescale enterprises are actually working against — not quarters, hours.
What works: governance built in before scale, not bolted on after
Two examples show the other side of this, and both come from regulated financial services, which is exactly where I do my own work.
A Singapore-headquartered global bank governs every AI use case — agentic included — through what it calls the PURE framework: purposeful, unsurprising, respectful, explainable. It’s been running since 2019, sits under a cross-functional Responsible AI Council with legal, risk, and technology at the table, and requires every use case to clear a materiality assessment before deployment. DBS Bank won Celent’s 2025 Model Risk Manager award for AI and GenAI off the back of it.
A 240-year-old U.S. custody bank takes governance to what I think is the most concrete level I’ve seen documented. BNY has built and deployed over 100 “digital employees”, each with a distinct persona, credentials, and a named human supervisor. Permissions are scoped tightly at the team level — no digital employee has broad cross-company access — and production changes still require human sign-off. Three separate oversight bodies sit above deployment: a data use review board, an AI release board, and an enterprise AI council. The bank’s own framing is worth repeating in substance rather than quote: good governance functions as a speed enabler, not a brake, because it lets the organization move faster with confidence instead of slower with doubt.
That’s the reframe I keep coming back to. Governance isn’t the tax you pay for scaling agentic AI. It’s the thing that makes scaling possible without finding out the hard way where the boundaries should have been.
The control stack I actually use: my Five-Level Agentic AI Roadmap
I didn’t wait for a standards body to hand me a checklist. I built one — the Agentic AI Roadmap, five levels, each earned by institutionalizing the one below it. It’s not a maturity score for show. It measures how reliably an organization turns intent into outcomes, and it’s the lens I run every deployment through, including the ones I architect in regulated origination.
| Level | Name | What’s actually true at this level |
|---|---|---|
| 1 | Prompted | No agents in production. People use copilots one at a time, value depends on who’s typing, and none of it is repeatable because none of it is captured. |
| 2 | Piloted | Task-specific agents run in bounded pilots with a human in the loop and a rollback path — but identity, logging, and scope differ pilot to pilot. Most enterprises are stuck here, which is why so many projects quietly get canceled. |
| 3 | Governed | Agents become an institutional asset instead of one engineer’s craft: a reference architecture, an identity standard, least-privilege scoping by default, and a registry so someone actually knows what’s running. |
| 4 | Assured | Behavior is measured, not asserted — decision quality, escaped-error rate, drift, cost per outcome — with an audit trail complete enough to satisfy an examiner. |
| 5 | Autonomous | The program improves itself. Guardian agents watch production agents, root-cause analysis feeds the reference architecture, and autonomy expands because control is provable, not assumed. |
The line that matters most out of all five: you don’t reach Level 5 by removing controls, you reach it by making control reliable enough that more autonomy becomes safe. An organization that jumps from Piloted straight to a marketed “autonomous” hasn’t followed the roadmap — it’s skipped the work that makes autonomy safe. That’s Deloitte’s 21% and Gartner’s 40%-plus cancellation number in one sentence.
If you’re stuck at Piloted, the fix isn’t a bigger pilot — it’s the reference architecture and the registry that get you to Governed. If you’re Governed, the fix is the measurement and audit trail that get you to Assured. Skipping straight to autonomy is how the cancellations happen — you can’t buy your way to Level 5, you have to earn Level 3 and 4 first.
Where I’d start if I were you
If your organization can’t currently name the owner of every production agent, that’s step one, before anything else on this list. Everything past that — permission scoping, monitoring, escalation design, audit trails — depends on knowing who’s accountable for what’s running.
I’d rather see an organization deploy fewer agents with real boundaries than chase Gartner’s 2028 number for its own sake. The 40%-plus project cancellation rate Gartner is already forecasting isn’t a technology failure. It’s what happens when the governance conversation starts after the agent is already in production.
Intent in, outcomes out — but only if someone’s watching the outcomes.
Go deeper with my newsletters:
- Agentic AI Architect — architecture and control design for enterprise agents
- Agentic AI Case Studies — deployment evidence and ROI teardowns
- Agentic AI P&L — cost, payback, and unit economics
- Agentic AI Governance — where autonomy meets accountability
© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026.
Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.
Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.
Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

