Agentic AI Statement of Work: Outcome, Not Effort

Agentic AI statement of work header reading "Specify the Outcome, Not

What is an agentic AI statement of work?

An agentic AI statement of work is the contract that says what an agent-run system must deliver and how acceptance is measured. It also sets the autonomy level and names who signs for each control. A traditional SOW buys effort, but this one buys an outcome with the evidence attached.

New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.

You cannot contract for effort when an agent does the work. You can contract for the outcome and the evidence.

Why does a traditional SOW break when agents do the work?

It breaks because it prices and governs effort, and effort is no longer what drives the cost. A classic SOW lists a staffing plan, a rate card and a set of documents. Then it pays for hours against milestones. That made sense when people wrote the code.

Now agents write most of it. So a team of five can ship what twenty used to ship, and the hours stop telling you much. The deliverables say even less. A design document proves that someone wrote a design document. It does not prove the agent will stay inside its limits at two in the morning.

The bigger gap is governance. Most SOWs I read still have no line for the autonomy level, the kill switch or who may expand either one. Those are the decisions that carry the risk. Yet the contract says nothing, so they default to the vendor.

What does an agentic AI statement of work specify instead?

An agentic AI statement of work specifies eight things that a traditional SOW leaves out or buries in an appendix. I built this list from my work in regulated loan origination, where a contract has to survive an audit, not just a kickoff.

Section Traditional SOW Agentic AI SOW
Scope Features and deliverables Intent, outcome metrics and the actions the agent may never take
People Roles, FTEs and hours Named Accountable roles, plus the vendor’s agent capacity
Acceptance UAT sign-off on test scripts Eval pass rate over a stated number of runs
Autonomy Not covered Autonomy level per workflow, human checkpoints, rules to move up
Controls A security appendix Tool permissions, credential scope, kill switch, trace retention
Change Change request on requirements Notice before any model, prompt, tool or policy change, with an eval re-run
Pricing Time and materials, or fixed bid Price per accepted outcome, plus capped token pass-through
Exit Code and documents Intent files, evals, prompts, skills and traces

Look at the right-hand column as a whole. Every row moves from describing work to describing a result and its proof. Your counsel will choose the words. My job as the architect is to make sure the rows exist at all.

How do you write acceptance criteria for a non-deterministic system?

You write them as a pass rate over many runs, never a single pass. An agent can take a different route to the same answer each time, which I covered in my post on non-determinism. So one green UAT script proves very little.

Instead, the contract names a fixed test set, a run count and a threshold for each risk tier. For example, a low-risk drafting task might need 90 percent over 50 runs. A credit decision might need 99 percent, plus a weekly human sample review. The eval suite becomes the acceptance gate, and the client owns it.

One more clause earns its place here. The vendor should not grade its own homework. Either the client runs the evals or a third party does, and the results go into the record.


More on Agentic AI Accountability



Who owns the autonomy level in the contract?

The client owns it, and the agentic AI statement of work should say so in plain words. I now put an autonomy schedule into every SOW I review. It lists each workflow, the level it runs at, the human checkpoints and the evidence needed to move up.

Without that schedule, autonomy drifts. Anthropic’s study Measuring AI agent autonomy in practice found that new users ran about 20 percent of Claude Code sessions on full auto-approve, while they had fewer than 50 sessions. By 750 sessions, it was more than 40 percent. Nobody signed a change for that shift. It came from habit, and in a contract, habit is not a control.

So expansion needs the client’s Governance lead to sign, as in my RACI matrix for agent projects. The vendor can recommend a higher level. It cannot decide.

How should an agentic AI statement of work price the work?

An agentic AI statement of work should price the accepted outcome, with model usage on a separate, capped line. Per-outcome pricing is already live in customer support. Intercom’s Fin charges $0.99 per outcome, and it counts a resolution when the customer asks for no further help after Fin’s last answer.

That definition is the whole contract. Change “no further help” to “ticket closed” and the bill moves, even though the work did not. So I spend more time on the outcome definition than on the price.

Tokens need their own clause, because someone has to pay when an agent loops for an hour. I want a monthly cap and a pass-through at cost or a stated markup. Then I add a rule that a runaway loop above the cap is the vendor’s problem. My post on cost per completed task has the math behind that line.

What happens when the model changes mid-contract?

A model change is a scope change, and the agentic AI statement of work should treat it as one. Library upgrades rarely change what your system decides. Yet a new model version can change every decision it makes.

So the clause asks for written notice before any model, prompt, tool or policy change goes live. Then it asks for a full eval re-run against the acceptance thresholds, plus the right to roll back. I also want a versioned list of everything the agent touches. Honestly, that list is what saves you in an audit, because it shows which version made which call.

What should the client own when the contract ends?

The client should own the intent files, the evals, the prompts, the skills and the traces. Those are the real work product now. Code can be regenerated. The intent.md and the eval set cannot, because they hold the client’s judgment.

A transition plan that hands over only code and documents leaves the client locked in. So I write the exit list on day one, while everyone is still friendly.

Instructions in, results out was IT. Intent in, outcomes out is agentic AI. A contract built for the first world buys hours and hopes for results. An agentic AI statement of work buys the outcome and names the owner up front.

I wrote the long version of this operating model in Intent In, Outcomes Out and the autonomy side in Earned Autonomy.

Here is my question for you. Open your largest AI services SOW. Does it say who may raise the agent’s autonomy level? If not, that is your first redline.

Go deeper

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI governance and architecture for regulated enterprises.

Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

 


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading