What Agentic AI Actually Costs: The Complete P&L Breakdown

Split-panel graphic showing agentic AI costs: budgeted per model call versus metered per completed task, by Dr. Harish Kotadia, Ph.D.

The cost of agentic AI is not the price of a model call. It is every token, tool call, retry, verification pass, escalation, and minute of human review it takes to finish one completed unit of work — and in the twelve months ending December 2025, the price per million tokens fell by roughly half while tokens consumed grew 4.5 times. Halve the price, multiply the volume by 4.5, and the spend line lands at 2.25 times where it started. That is my arithmetic on two figures Bain published in June 2026, and it is the single most expensive misunderstanding in enterprise AI planning right now.

Falling token prices are what most planning models lean on. Bain’s finding is that they get absorbed by heavier usage per task. Unit price is not the lever. It never was.

Here is my working definition, and why the budget breaks: agentic AI is the next enterprise workload abstraction — a governed, goal-driven software layer where LLM-powered agents autonomously plan and execute multi-step business processes. For a CFO the operative word is autonomously. Software that plans its own steps decides its own consumption. A cost you cannot forecast from headcount or seat count is a budget problem before it is a control problem.

The six cost lines every deployment carries

Every agentic AI deployment I have reviewed carries the same six lines — I first mapped them in The Six-Line Stack. Most companies budget for one.

  1. Inference and model spend — variable, uncapped. In one disclosed rollout at a global ride-hailing company, per-engineer monthly cost averaged $150 to $250, with heavy users between $500 and $2,000. Bain reports the top 5% of users inside a company frequently consume more tokens than the remaining 95% combined.
  2. Orchestration and infrastructure — semi-fixed. Undisclosed in every case I reviewed.
  3. Human review and escalation — variable, at loaded labor rates. Undisclosed.
  4. Integration and build — one-time, usually capitalized. Trade estimates cluster between $150,000 and $800,000, but those are vendor figures, not audited ones.
  5. Evaluation, monitoring, and audit evidence — fixed and permanent, and the line most often missing entirely from the original business case.
  6. Rework and failure cost — variable. Nobody forecasts it and nobody discloses it.

One line quantified out of six. That ratio is the honest state of public disclosure on agentic AI economics, and any ROI claim built on the other five is a model, not a finding.

Capacity is not cash

Separate the two, because a controller will. That ride-hailing company reported roughly 70% of committed code originating from AI and 11% of live backend updates written by agents with no human in the loop. Production numbers, not pilot numbers. They are also capacity, not cash: its president said in May that the line from those statistics to shipped consumer features could not yet be drawn. Capacity converts to cash only when headcount, backfill, contractor spend, or overtime changes. None had, and hiring continued.

So the correct entry against a 2.25x cost line is zero realized savings and an unquantified productivity gain. A CFO will accept that entry. A CFO will not accept it being called ROI.

The month-four blowup

The same company exhausted its entire calendar-2026 AI budget by April — four months in — against 2025 R&D spend of $3.4 billion. Adoption of one agentic coding tool went from 32% to 84% of a 5,000-engineer organization inside a quarter. One engineer ran up $1,200 in a single two-hour session. The chief technology officer told The Information he was rebuilding his assumptions from scratch.

Nothing failed. The tool did what it was sold to do, the engineers used it as intended, and the forecast was still wrong by a full year. That is the risk worth costing: not a failed deployment, a correct deployment on the wrong budget line.

On my Five-Stage Agentic AI Roadmap — Prompted, Piloted, Governed, Assured, Autonomous — consumption at 11% unsupervised production commits is Autonomous-stage behavior. The budget was drawn at Piloted, where spend is discretionary, annual, and small enough to sit inside an existing line item. Late-stage consumption funded on an early-stage budget line is the configuration that blows up, and it blows up in month four, not month twelve.

What actually moves the economics

Model routing is the lever. A US telecom carrier running roughly 8 billion tokens a day reorchestrated so that large coordinating agents route work to smaller domain-specific models rather than sending everything to frontier. It reported a 90% cost reduction with three times the throughput. One architectural choice moved the economics by an order of magnitude.

Incentives are a cost driver too. The ride-hailing company ranked engineers on internal leaderboards by tool usage. Paying per token while rewarding consumption is a control failure with a dollar value.

And the correction is documented — it was not a spend cap. On 5 August the same CTO said cost per token had fallen even as frontier-tool users quadrupled, achieved through better prompt caching, a cheaper default model, efficiency evaluation of new models, and giving engineers hourly visibility of their own spend. His framing, reported by Fortune: they treated efficiency as an engineering problem rather than a budget problem. I would go further. Three of those four levers are configuration, and configuration is auditable. That makes token efficiency a testable control, not a good intention.

The 30x nobody budgets: the agent checking its own work

A customer-service interaction that cost $0.04 in 2023 costs about $1.20 in 2026 — a 30x unit-cost increase as the same interaction moved from a single linear model call to an orchestrated multi-step agent (EY, June 2026, a consultant-modeled comparison). EY’s own note is that the figure excludes evaluation and human-collaboration design, so treat $1.20 as a floor.

The line driving it is response refinement: the iterative checking, correcting, and improving an agent runs internally before showing a human anything. It is inference spend, priced at the same per-token rate as the response, but it buys verification rather than output. McKinsey’s July 2026 analysis puts it at 60% of agentic operating cost — more than the first inference call and everything else combined. Its Enterprise AI FinOps Survey found 93% of respondents have already exceeded their AI budget; its 2026 State of AI survey found one in five organizations have actively constrained AI use over operating cost.

The instinct is to treat this as a model-price problem and negotiate a cheaper rate card. That compounds rather than corrects: a 50% cheaper model still spends 60% of its smaller bill checking itself. Only the refinement rate moves the ratio — and cutting verification passes cuts cost, yes, along with the evidence that any given output was checked.

The return-side data tells the same story from the other direction. A 2024 Microsoft-sponsored IDC study found $3.70 back per $1 spent on generative AI, top performers near $10.30 — but only about 6% reach that tier. Deloitte’s 2026 survey of 3,235 leaders across 24 countries found 66% report productivity gains and only 20% see AI-driven revenue growth. Usage and spend are up. The share converting into verified, monetizable output is the bottleneck, not the sticker price of a token.

Budgeting the refinement line differently doesn’t fix it. Governing it does, and that is what I built the Earned Autonomy Index to score: a 0–1000 rating across five pillars, one of which — Outcome Assurance — asks whether degradation in the agent’s output would be noticed at all. Applied to 507 publicly documented enterprise deployments, three-quarters sit in the Piloted band and not one reaches Autonomous; Outcome Assurance, together with Intent Specification, caps 470 of the 507. The field has built the controls that stop an agent doing damage, not the ones that would tell anyone it had quietly stopped doing its job well. The full methodology is in my book Earned Autonomy; the cost-stack and roadmap framing is in Intent In, Outcomes Out.

Why per-hour pricing had to die

Every pricing model is a bet on what’s scarce. Time-and-materials bet that hours were scarce. Per-seat bet that logins were scarce. Fixed-bid bet that scope was the thing worth pricing. Agentic AI breaks all three bets at once, because the scarce resource stopped being human time. It’s compute now, and compute doesn’t clock in, doesn’t get tired, and definitely doesn’t stay inside the estimate you gave the client in March.

What replaces it — price per decision — is not one number. It’s four layers, and the mistake I keep seeing vendors make is pricing one or two of them and calling it done.

Layer 1 — Base or platform fee. Standing up the agent: integration, source-system connections, initial governance configuration. The one piece that still looks like the old world.

Layer 2 — Consumption unit. Priced against the real variable driver: tokens, API calls, agent actions. A five-step reasoning task can eat five to thirty times the tokens of a simple chatbot query for what looks, on paper, like the same request — that’s Gartner’s number, from March 2026. A rate card that can’t tell those two requests apart isn’t pricing anything. It’s guessing.

Layer 3 — Outcome unit. The layer buyers actually came to buy: price per resolved case, per verified decision, per completed workflow. Intercom’s Fin charges $0.99 a resolution; Sierra runs pure outcome pricing with third parties pegging it around $1.50 per resolved interaction. It’s the layer that feels like value — and exactly the layer most likely to get quietly gamed if “resolved” turns out to mean the customer went silent rather than the problem got solved.

Layer 4 — Risk and assurance premium. Nobody’s pricing this one yet. I’d argue it matters more than the other three combined. It covers the audit trail, the governance controls, the escalation coverage — everything that lets you prove a decision landed inside its assigned boundary instead of trusting that it did. Skip it and outcome pricing becomes unverifiable by design: you’re paying per resolution on the vendor’s word for what counted as resolved. That’s not a pricing model. That’s an honor system with an invoice attached.

The vendors getting outcome pricing right treat Layer 4 as something they sell, not overhead they eat. Which is exactly the gap Gartner is pointing at when it projects more than 40% of agentic AI projects canceled by the end of 2027 — not because the models don’t work, but because the governance and scoping around them never got built. Build the risk and assurance layer first. Quote a price per decision second.

What a risk committee would still ask

Who holds the budget authority to stop consumption mid-quarter, and has that authority ever been exercised? What share of spend is refinement versus first response — and if that split isn’t measured, on what basis is the current budget defensible? If cost per task is not instrumented today, on what basis was this approved?

The last one is uncomfortable, and in my work in regulated loan origination it is the one that lands. Cost per decision is the same metric as cost per exception. An institution that can price a manual underwriting decision to the cent and cannot price an agentic one has not finished the business case.

Would most of these deployments survive a budget review as originally submitted? No. They’d survive the resubmission — the version with a metered unit, an owner, a routing policy, and a visible run-rate. Nothing about the technology changes between the two. Only the instrumentation does.

What does one agentic task cost you today, and can you prove it?

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.

Disclaimer: This blog post is based on recent events and news items drawn from reputed media sources and vendor websites available in the public domain and quoted above. This post is intended for educational purposes, to help the enterprise agentic AI community learn from public information on the application and use of agentic AI tools and technology in Fortune 500 companies.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of published news reports quoted above and does not constitute legal, financial, or consulting advice of any kind.


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading