Agentic AI Definition of Done: Done Is a Pass Rate

Agentic AI definition of done header reading "Done Is a Pass Rate," with three Thinkers360 Certified Expert badges and a Three Anonymized Case Studies call-out, by Dr. Harish Kotadia, Ph.D.

What is agentic AI definition of done?

An agentic AI definition of done is the measured bar an agent must clear before anyone calls the work finished. Not “the demo worked.” Not “the tests went green once.” A pass rate on real tasks, scored the same way every run, with one named person who signs for it.
New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.
You cannot declare an agent done. You can measure when it is.

Why does “done” break for agents?

Because the old definition assumed the code does the same thing twice. In traditional IT, done meant the feature shipped and the tests passed. Run them again tomorrow and they still pass. An agent does not give you that. The same prompt on the same input can take a different path on Tuesday than it took on Monday. So one green run tells you the agent can work, not that it does.

I learned this the slow way in my work in regulated loan origination. An agent cleared a 40-case review on a Thursday. The next week it cleared 36 of the same 40. Nothing had changed except the dice. Since then, my agentic AI definition of done has carried a number.

What happens when “done” means “it answered”?

Two public cases show the bill. Neither team had a bad model. Both had no bar.

A national airline put a customer service chatbot on its website. The bot told a grieving customer he could claim a bereavement fare after he flew. The real policy said the opposite. Done had meant “the bot answers,” and nobody had scored its policy answers against anything.

The second case is closer to home for anyone shipping agents. A software founder ran a public twelve-day build on an agentic coding platform. He judged progress on demos and on the agent’s own status reports. A “code freeze” existed as one sentence in a prompt. On day nine the agent deleted the production database during that freeze, wiping records on about 1,200 executives, then told the founder a rollback was impossible. It was not. The platform’s CEO called the deletion “unacceptable” and shipped dev and prod separation within days.

Both stories end the same way. Someone outside the team, a tribunal or a founder on social media, did the acceptance test the team never wrote. That is what a missing agentic AI definition of done looks like in practice.

What does a measured “done” look like?

A large online travel marketplace shows the other outcome. The team needed to move nearly 3,500 test files from one testing framework to another. The manual estimate was a year and a half. So they built an agent loop and gave it a per-file bar. The rewrite had to pass the test runner, then lint, then the type checker. Only then did the file move to complete. No opinion in the loop. The file cleared the gate or it went back for another try.

The first bulk run took four hours and cleared 75 percent of files. Four days of tuning on failing samples pushed that to 97 percent. The team retried the hardest files up to 100 times, then stopped the automation and finished the last 3 percent by hand. Six weeks, end to end.

What I like most is that nobody argued about done. The gate decided. That is the agentic AI definition of done I want on every project. A machine-checked bar per unit of work, a pass rate someone reports, and a human who decides when to stop automating.

Question Traditional SDLC Agentic AI
What proves it works? Tests pass once A pass rate across repeated runs
Who decides? The author and a reviewer An eval suite plus one named signer
What is the unit? A feature or a ticket A task class with a sample of real cases
When do you stop? All tests green The pass rate plateaus; a human finishes the rest
Does done stay done? Yes, until the code changes No; a regression suite re-checks it after every change

More on Agentic AI SDLC



How high should the bar be?

High enough that the agent clears it every time, not on average. Anthropic’s engineering team says regression evals “should have a nearly 100% pass rate.” For customer-facing agents they recommend pass^k, which asks whether all k trials succeed. Their arithmetic is sobering: an agent with a 75 percent per-trial success rate, run three times, passes all three only about 42 percent of the time.

So I set two numbers. Regression sits at near 100 percent, because a task the agent handled last month has to stay handled. Capability starts wherever the agent is today and rises each release. A global wealth manager runs the same shape: it tests every AI use case before deployment and re-runs a regression suite of sample questions daily. That is a regulated firm’s agentic AI definition of done, even if it never used the phrase.

What goes into my definition of done?

Five things go into my agentic AI definition of done, and none of them is a demo.

  • A task set: 20 to 50 real cases drawn from real failures, each with a pass or fail answer two experts would agree on.
  • A pass rate, not a pass: The agent runs the set several times; the number I sign is the worst run, not the best.
  • A regression floor: Near 100 percent on everything it already handled, checked after every prompt, model or tool change.
  • A named signer: One person owns the number, and it is never the person who built the agent.
  • A stop rule: When the pass rate plateaus, a human finishes the tail instead of the agent retrying forever.

Instructions in, results out was IT. Intent in, outcomes out is agentic AI. An outcome without a number attached is only a hope. So the number I sign is a pass rate, and it is written down before the build starts.

I wrote the long version of this operating model in Intent In, Outcomes Out and the autonomy side in Earned Autonomy.

What number does your team sign before it calls an agent done?

Go deeper

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI governance and architecture for regulated enterprises.

Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

 


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Leave a Reply

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading