Agentic AI Lesson: Governed Scales, Headcount-First Gets Rolled Back

Two enterprise agentic AI banking case studies compared — a governed custody bank deployment that scaled versus a headcount-first voice bot rollback

Agentic AI is software that takes a goal, plans the steps, uses tools and takes action toward an outcome with limited human supervision — not a chatbot that only answers questions. I picked two 2025–2026 banking deployments from public reporting because they ran the same play in the same industry and landed in opposite places. One scaled agents into production and reported measurable gains. One cut the humans first, and had to hire them back within weeks.

The difference was not the model. It was whether the organization built the governance before it claimed the outcome.

What Is Agentic AI, and Why the Roadmap Matters Here

Agentic AI moves the enterprise from “instructions in, results out” to “intent in, outcomes out.” I define it fully in my pillar post on what agentic AI is. The short version: an agent reasons, plans, calls tools and APIs, and acts — and that capability is exactly what makes governance non-optional.

I map every deployment I analyze to my Five-Stage Agentic AI Roadmap: Prompted → Piloted → Governed → Assured → Autonomous. The levels are cumulative. You do not reach autonomy by removing controls; you earn it by making control so reliable that more autonomy becomes safe. This is the same lens I used in my earlier one-contained-one-catastrophic comparison.

The Case Studies: One Success, One Failure

Both organizations are banks. Both deployed AI agents into real work. One operated at Governed on my Roadmap. One operated at Piloted while marketing the outcome of Autonomous.

Source: case details below are drawn from public reporting, accessed July 2026. All figures and outcomes are as reported in the cited coverage.

Case Study 1 — The Success: A 240-Year-Old Global Custody Bank

A 240-year-old global custody bank
Industry: Banking / Asset Servicing

What they built: An internal multi-agent platform where 20,000+ employees build their own task-scoped agents, plus a fleet of 134 “digital employees” — agents with their own login, a narrow task scope, and a human manager — running repetitive operational workflows such as validating payment instructions overnight.

Stack: A governed internal AI platform connecting multiple frontier LLMs (including OpenAI models) to the bank’s data and compliance systems, with permissions, security and oversight standardized at the platform level.

Outcome: A contract-review assistant cut legal review time by 75%, from four hours to one, across 3,000+ annual vendor agreements. Roughly 99% of the workforce was trained on the platform. The bank spent about $3.8 billion on technology in 2025, ~19% of revenue — the highest proportion among its peers, per CNBC.

What makes this a Governed (Stage 3) deployment, not a Piloted one: governance lives inside the tooling. As the bank’s deputy general counsel put it, the platform “embeds governance at the system level… standardizes permissions, security, and oversight across all models and tools.” The agents augment people rather than replace them. The head of payment operations told CNBC the digital employee “works 24/7… focused on very specific repetitive tasks that allow our human employees to do much more human, intense, interesting-type roles.” The human’s job shifted to training and supervising the agent — intent in, outcomes out, with accountability attached.

Case Study 2 — The Failure: Australia’s Largest Bank

Australia’s largest bank
Industry: Banking / Retail Customer Service

What they built: An AI “voice bot” deployed to a direct-banking call center to automate simpler customer queries, positioned as reducing call volumes enough to make 45 human customer-service roles redundant.

Stack: A generative-AI voice agent fronting the contact center, deployed alongside a workforce reduction rather than behind a bounded pilot with defined success metrics.

Outcome: The bank said the voice bot reduced call volumes by 2,000 a week but the volumes were in fact rising, forcing the bank to offer staff overtime and direct team leaders to answer calls. The bank reversed the 45 redundancies on August 21, 2025.

This is a Piloted (Stage 2) capability sold as an Autonomous (Stage 5) outcome. The tell is the metric: the headcount decision was made on a claimed reduction in call volume that the deployment could not actually sustain. There was no assured measurement layer — no defensible, examiner-ready evidence that the agent was in control of the workload before humans were removed from it. When reality diverged from the slide, there was no graceful degradation, only managers back on the phones and a public reversal within weeks.

The Lesson: Governance Comes Before the Outcome You Claim

Put the two side by side and the pattern is unmistakable.

  • The custody bank earned its autonomy. Governance was enforced by the platform — identity, least-privilege scope, oversight — so scaling agents was safe, and the outcomes (75% faster contract review, 24/7 operations) were measured, not asserted.
  • The retail bank claimed the outcome before it earned the assurance. It removed the humans on the strength of a productivity story the system couldn’t back with data, and paid for it in overtime and a reversal within weeks.

This is the “intent in, outcomes out” turn from my Roadmap. It only holds if your organization can make that turn reliably and account for it afterward. Autonomy is not a property of the model — it is a property of the controls you can prove.

If you are deploying agents into customer-facing or regulated work, the question to ask before go-live — not after — is simple: is the outcome you are about to announce backed by enforced controls and measured evidence, or by a prompt and a hope?

Conclusion

Governed agentic AI scales. Headcount-first agentic AI gets rolled back. Map your deployment to the Roadmap honestly, earn assurance before you claim autonomy, and never cut the humans before the evidence says you can.

Want to know which Roadmap stage your agents are really on? Start with the Agentic AI Roadmap and place yourself by your weakest production agent, not your best pilot.

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

 

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.

He is the author of Intent In, Outcomes Out: A Practitioner’s Field Guide to Agentic AI for the Enterprise, available on Amazon at: https://mybook.to/AgenticAI

He writes the Agentic AI Architect newsletter on LinkedIn: https://www.linkedin.com/newsletters/agentic-ai-architect-7477724166675140608/

What is Agentic AI? Agentic AI is a governed, goal-driven software layer in which LLM-powered agents autonomously plan and execute multi-step business processes. Instructions in, results out, was IT. Intent in, outcomes out, is agentic AI. Read the full definition at: https://agenticaiarch.com/what-is-agentic-ai/

The Agentic AI Roadmap is a five-stage maturity model — Prompted, Piloted, Governed, Assured, Autonomous — measuring how reliably and accountably an organization turns intent into outcomes. Explore the roadmap at:
https://agenticaiarch.com/agentic-ai-roadmap/

Follow him on LinkedIn at www.linkedin.com/in/hkotadia and at www.agenticaiarch.com


Discover more from Agentic AI Architecture | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Architecture | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Architecture | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading