Agentic AI does not cost what your IT budget assumes it costs. Traditional IT financial management treats compute as a fixed, forecastable line item — CapEx depreciated over years, or OpEx metered but still tied to a workload you can define in advance. Agentic AI breaks that assumption at the root: cost now scales with autonomy and reasoning depth, one decision at a time, and it is driven by behavior your procurement team never modeled.
That single habit — modeling compute as infrastructure spend instead of per-decision spend — is one of the ten IT-era instincts I believe practitioners need to unlearn to succeed in agentic AI. This post goes deep on that one, because for most enterprises it is the one that shows up on the P&L first.
The Definition: What “Per-Decision Cost” Actually Means
Per-decision cost is the total spend an agentic system incurs to complete one unit of autonomous work — a resolved ticket, a completed research task, a processed loan exception — measured across every token, tool call, retry, and orchestration step involved, not just the sticker price of a single model call.
This is a different unit of account than IT has ever budgeted against. A cloud VM has a knowable hourly rate. A SaaS seat has a knowable monthly rate. An agent’s cost is emergent: it depends on how many steps the model takes to reason through a task, how many tools it calls, how many times it retries, and how many other agents it coordinates with. According to a Gartner press release on agentic AI coding costs, agentic workloads consume 5 to 30 times the tokens per task of a standard chatbot exchange. Anthropic, describing in its own engineering blog how it built its multi-agent research system, reports that “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats” — with a single research session capable of burning through millions of tokens and dollars per query.
That is the mechanism behind the paradox every CFO needs to understand: per-token prices are collapsing while enterprise AI bills are exploding. The Stanford AI Index documented inference cost for a GPT-3.5-class system falling more than 280-fold in two years. Menlo Ventures tracked enterprise generative-AI spend rising from $2.3B in 2023 to $37B in 2025. Cheaper tokens, more expensive systems — economists call this the Jevons Paradox, and it is exactly why “compute is a fixed line item” stops working the moment you deploy an agent instead of a chatbot.
Why Traditional IT Financial Management Wasn’t Built for This
FinOps — the discipline that governs cloud spend through visibility, allocation, and optimization — was a real advance over CapEx budgeting. But even FinOps assumes a workload procurement can define. As NPI Financial explains in No Jitter, AI spend “is consumption-driven, but unlike cloud infrastructure, it is not tied to a workload that procurement can define or forecast in advance. It is tied to user behavior.” In that same piece, Gartner analyst goes further: the sources driving AI spend now sit “even outside of our organization” — end users and customers, through how they prompt applications, directly determine the enterprise’s bill.
That is the gap. Here is what it looks like side by side:
| Traditional IT Cost Model | Agentic AI Cost Model | |
|---|---|---|
| Unit of account | Server, seat, license | Token, tool call, resolved task |
| Predictability | Forecastable from provisioned capacity | Emergent from reasoning depth and autonomy |
| Who drives cost | IT provisions it | End-user and agent behavior drives it |
| Budget cadence | Annual, reviewed at year-end | Continuous, monitored in near real time |
| Governance lever | Capacity planning | Runtime circuit breakers, spend ceilings |
| Failure mode | Over-provisioning (waste) | Runaway loops (blowout) |
EY’s “Total Cost of Agents” series makes the same point in a single line: “the total cost of an agent, at most organisations, is structurally invisible until designed for visibility.” EY’s model breaks that cost into seven line items — infrastructure, governance, organizational change, failure recovery, and emerging regulatory risk among them — and notes that “most companies only include costs 1-3 in their agentic investments and business cases, with costs 4-7 often emerging later in the lifecycle as agents scale.” EY’s framing of the fix: treat agents “as growth investments, with clear ownership, cost controls and value metrics.”
The consequence, per a Gartner poll cited in that same EY analysis, is stark: more than 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. Separately, Gartner’s own newsroom coverage projects that at least 50% of GenAI projects will overrun their budgeted costs through 2028. Those are predictions, not certainties — but they are the predictions of the analyst firm enterprises pay to get this right.
The Roadmap: Moving from Fixed Spend to Governed Per-Decision Cost
I break this into four stages. This mirrors the assurance discipline I laid out for agent autonomy itself in Earned Autonomy — the same principle that governs whether an agent has earned the right to act without a human in the loop also governs whether an agent has earned the right to spend without a ceiling. Autonomy and cost are not separate risks; they are the same risk measured on two different axis of my EAI-1000 framework.
- Establish visibility (0–90 days). Turn on per-call cost logging — model, token counts, and the workflow that triggered the call — for every agent in production. You cannot govern what you cannot see, and most enterprises today cannot answer a simple question: what does one unit of agentic work cost?
- Install controls before you scale. Before your agent portfolio grows past three or four use cases, put runtime circuit breakers in place — hard budget caps per run, per hour, per day; loop and step limits; spend-velocity throttling; and human-approval gates before irreversible or high-cost actions. BaristaLabs’ spend-circuit-breaker framework makes the key point directly: these controls have to live outside the agent’s own code, enforced at a policy layer the agent cannot talk its way around. Finout’s agentic cost-governance research adds that a monthly cap alone isn’t enough — “a $500/month cap still allows burning $500 in 20 minutes” without spend-velocity monitoring layered on top.
- Institutionalize Agent FinOps. Name an owner — a Head of Agent Economics — accountable for the full cost picture, not just the model bill. Require a value metric for every agent before it goes live: output, revenue, or risk reduction per dollar spent, not just capability demonstrated in a demo.
- Renegotiate the contracts. As a buyer, demand unit-cost transparency, overage definitions, and consumption caps from every AI vendor. As a services provider, recognize that outcome-based and consumption-based pricing are displacing seat-based and time-and-materials pricing — and that the firms that move first on unit economics will set the terms for everyone who moves second.
For a deeper walkthrough of the maturity stages behind that roadmap and how they map to my five-stage agentic AI roadmap, see my companion post, What Is Agentic AI?, and the documented case evidence in the Agentic AI Case Study Rolodex, where outcome assurance — the same discipline that catches cost risk — is the capping factor in just over half of the 507 cases I’ve catalogued.
What This Means for Clients and IT Services Firms
For enterprise clients, the flat-rate SaaS mental model — “how many seats do we need?” — stops being the right question. The right question becomes: what behaviors drive our usage, and who owns the budget when that usage spikes without warning? Procurement teams need new contract language — unit definitions, overage pricing, consumption ceilings — before they need new vendors.
For IT services and consulting firms, the exposure is structural, not incremental. Four decades of an industry built on growing revenue by growing headcount is colliding with a delivery model where cost and value no longer scale with people — they scale with decisions. Gartner analyst has described this directly: AI-driven automation is “shifting the competitive advantage from labor arbitrage to technology arbitrage,” forcing a reassessment of margin structure and accelerating contract renegotiation across managed services. The firms that build per-decision cost governance into their own delivery model — rather than discovering it after a client’s invoice triples — are the ones that will price this transition on their own terms instead of having it priced for them.
I’ll go deeper on the pricing-model side of this shift — outcome-based versus consumption-based contracts, and what it means for T&M billing specifically — in an upcoming edition of my Agentic AI P&L newsletter.
Intent in, outcomes out was never just a phrase about what agentic AI does. It’s a phrase about what agentic AI costs. Every outcome an agent produces was purchased one decision at a time — and until you can measure that purchase, you cannot govern it, price it or defend it on your P&L.
Related reading on my site: my definition of agentic AI, my five-stage agentic AI roadmap, and my library of enterprise agentic AI case studies. My governance methodology in full is in Earned Autonomy, available on Kindle.
© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026.
Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.
Disclaimer: This blog post is based on recent events and news items drawn from reputed media sources and vendor websites available in the public domain and quoted above. This post is intended for educational purposes, to help the enterprise agentic AI community learn from public information on the application and use of agentic AI tools and technology in Fortune 500 companies.
Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of published news reports quoted above and does not constitute legal, financial, or consulting advice of any kind.

