The Difference Was Operational Risk Management
Summary:
Of the two 2025 agentic AI coding deployments compared here, one sits at Stage 3, Governed, on my Agentic AI Roadmap, and one operated at Stage 2, Piloted, while being marketed as Stage 5, Autonomous. The gap between them was not model capability. It was whether the agent’s permissions were enforced by architecture or requested by prompt.
I selected these two from public reporting on enterprise and consumer coding agent deployments and mapped each to my Five-Stage Agentic AI Roadmap, the same way I mapped the healthcare vs. fintech autonomy trap comparison. Below, each case states what was built, the permission model underneath it, and the reported outcome.
Source: case details below are drawn from public reporting on both deployments, accessed July 2026. All figures and outcomes are as reported in the cited coverage.
Where they land on the Agentic AI Roadmap
I plotted both against the five levels of the Agentic AI Roadmap, permission model against outcome.
| Deployment | Roadmap level claimed | Roadmap level evidenced | Outcome |
|---|---|---|---|
| Wall Street investment bank, coding agent
|
Governed | Governed | 3-4x productivity gain, zero uncontrolled incidents reported |
| Consumer coding platform, coding agent
|
Autonomous | Piloted, at best | Production database deleted, records fabricated afterward |
What Went Right
A marquee Wall Street investment bank deployed an autonomous AI software engineer alongside its twelve-thousand-strong developer workforce, starting with hundreds of agent instances on tasks developers consider drudgery, like migrating internal code to newer programming languages. The agent never held standing access to production. Every task arrived through the same ticketing workflow a human engineer would use, and every output routed through a supervising engineer before anything shipped. The bank’s technology chief was explicit that the engineer’s job now includes describing problems clearly and supervising the agents doing the work. Internal estimates point to a three-to-four-times productivity improvement over prior AI tooling, with instances scaling only as use cases proved out.
What Went Wrong
A consumer AI-powered app development platform learned the opposite lesson in public. A well-known SaaS investor spent twelve days building a commercial-grade app using the platform’s coding agent. On day nine, despite an explicit code freeze and repeated instructions not to make changes, the agent ran destructive commands against the live production database and deleted records for more than a thousand executives and companies. It then fabricated data and generated misleading status messages, including a false claim that rollback was impossible. The platform’s own remediation, shipped within days, tells you where the architecture failed: automatic separation between development and production databases, improved rollback, and a planning-only mode. The agent had held write access to production the entire time, and the only thing standing between it and the database was a natural-language instruction.
The most instructive pattern here isn’t the incident itself, it’s the fabrication that followed it. Once an agent can mutate production, the same capability lets it mutate the account of what it did, which is exactly why the bank’s supervising-engineer review gate and the platform’s prompt-only warning are not different degrees of the same control. They are different categories of control. One is enforced by the system. One is a request the system has no way to hold the agent to.
Autonomy is not a property of the model. It is a property of the permissions granted to it. That is why governed comes before goal-driven in my definition of agentic AI, and why Stage 3 comes before Stage 5 on the Roadmap. If an agent can write, delete, or move anything of value, the question to ask before go-live, not after, is what the maximum damage looks like on its worst day, and whether that ceiling is enforced by architecture or by instructions.
© Dr. Harish Kotadia, Ph.D. All Rights Reserved, 2026
Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.
Disclaimer: This blog post is based on recent events and news items drawn from reputed media sources and vendor websites available in the public domain and quoted above. This post is intended for educational purposes, to help the enterprise agentic AI community learn from public information on the application and use of agentic AI tools and technology in Fortune 500 companies. Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of published news reports quoted above and does not constitute legal, financial, or consulting advice of any kind.

