Agentic AI SDLC Best Practices: Close the Loop Before You Leave

Agentic AI SDLC best practices header in a white newspaper layout with the headline Close the Loop Before You Leave, by Dr. Harish Kotadia, Ph.D.

 What are agentic AI SDLC best practices?

Agentic AI SDLC best practices are the working rules for a lifecycle where an agent reads the code, plans, edits, runs commands and checks its own work. A human watches, redirects or walks away. Most of them come down to one constraint: the agent’s context window fills fast, and its judgment degrades as it fills. So the practices that hold up keep the loop short.

You cannot watch every turn. You can give the agent a check it can run. 

New here? I publish one agentic AI governance post every weekday. Subscribe to the blog and it lands in your inbox the moment it goes live.

Why does context management come first?

Because almost every other practice follows from it. The context window holds every message, every file the agent read and every command it ran. A single debugging session can burn tens of thousands of tokens. Then the agent starts forgetting earlier instructions and making more mistakes. The fix is to treat context as the scarcest resource in the lifecycle. Start a fresh session between unrelated tasks. Send research to a sub-agent with its own window so the main thread stays clean. And after two failed corrections on the same issue, stop correcting. Clear the session and write a better first prompt with what you learned. I have seen this rule save an afternoon more than once.

What does “give the agent a way to verify its work” mean?

It means the agent stops when the work looks done, so “looks done” must not be the only signal. Give it a test suite, a build exit code, a linter, a diff against a fixture or a screenshot compared to a design. Then the loop closes on its own: do the work, run the check, read the result, iterate until it passes. The guide is blunt about the alternative. Without a check, you become the verification loop. Every mistake waits for you to notice it. That is not a productivity note. It is the difference between a session you supervise and one you can leave. In my work in regulated loan origination, the check is the control. An agent that cannot prove the build passed is an agent that did not finish. I also ask for evidence, not assertion: the test output, the command it ran, what came back. Reviewing evidence is faster than re-running it, and it works for the sessions I never watched.

How should the work be sequenced?

Explore first, then plan, then code, then commit. The guide recommends a read-only planning mode for the first two phases. The agent reads files and answers questions but changes nothing. You review the plan, edit it if you need to, and only then let it implement against that plan with tests. Two caveats matter. Planning adds overhead, so skip it when you could describe the diff in one sentence. And for larger features, have the agent interview you first and write a spec. Then start a fresh session to execute that spec. The clean session has nothing in it but the work.

What belongs in the project context file?

Every platform has one now: a short file the agent reads at the start of every session. The guide’s rule for it is the best line in the document. For each line, ask whether removing it would cause the agent to make mistakes. If not, cut it. A bloated context file causes the agent to ignore your actual instructions. Include the commands it cannot guess, the style rules that differ from defaults, the repo etiquette and the environment quirks. Exclude anything it can read from the code, standard conventions, long tutorials and anything that changes often. Treat the file like code. Check it in, prune it, and test changes by watching whether behavior actually shifts.


More on the Agentic AI SDLC



Where do hooks, skills and sub-agents fit?

They are three different kinds of control, and agentic AI SDLC best practices only work when teams keep them apart. A context file instruction is advisory. A hook runs a script at a fixed point in the workflow, every time, with zero exceptions. The guide says it plainly: use hooks for actions that must happen every time. Lint after every edit. Block writes to the migrations folder. That is enforcement, not a suggestion. Skills hold domain knowledge and repeat workflows the agent loads on demand. So the always-on context file stays short. Sub-agents run in their own context with their own allowed tools. The guide uses them for two jobs I care about: research that would otherwise flood the main window, and hostile review. A reviewer in a fresh context sees only the diff and the criteria, which is why I wrote that one agent cannot check itself. It does not see the reasoning that produced the change, so it cannot be charmed by it. One warning from the guide belongs on the wall. A reviewer asked to find gaps will find some, even when the work is sound. Tell it to flag only gaps that break correctness or the stated requirements. Otherwise you get defensive code and tests for cases that cannot happen.

How do you scale from one session to many?

Run the agent from scripts and CI with no one at the keyboard. Run parallel sessions in isolated worktrees so edits do not collide. Fan a migration out across a file list, one invocation per file, with the allowed tools scoped tight because nobody is watching. Test the prompt on two or three files before you run it on two thousand. The guide also describes a writer and reviewer pattern across two sessions. The same idea works with tests: one session writes them, another writes the code to pass them. A fresh context reviews better because it is not biased toward code it just wrote. Anthropic’s own autonomy measurement found that new users auto-approve about one action in five. By their 750th session that rate has roughly doubled. Scale arrives whether you designed for it or not. So design the checks first.

What do the failure patterns have in common?

The guide closes with five. I put them in a table because each one is a governance failure with a technical name. So agentic AI SDLC best practices and governance controls turn out to be the same list.

Failure pattern What went wrong The fix
The kitchen sink session Unrelated tasks share one context Start a fresh session between tasks
Correcting over and over Failed approaches pollute the window After two corrections, restart with a better prompt
The over-specified context file Important rules get lost in noise Prune, or convert the rule to a hook
The trust-then-verify gap Plausible code with no check If you cannot verify it, do not ship it
The infinite exploration Unscoped research fills context Scope it, or send it to a sub-agent

Read the third column again. Every fix is a control on the harness, not a change to the model. That is the argument I have made on this blog for months. 

What do I take from this?

Three things, and I would apply them on any platform.

  • First, agentic AI SDLC best practices start with a runnable check, because the check is what lets a human leave the room.
  • Second, the context file is policy, and policy that nobody prunes stops being read.
  • Third, hooks and fresh-context reviewers are where the guidance stops being advice and starts being enforcement.

Instructions in, results out was IT. Intent in, outcomes out is agentic AI. The agentic AI SDLC best practices in this guide are how you write the intent down and prove the outcome. If your team has adopted a coding agent, which of these five failure patterns is costing you the most time right now?

Intent In, Outcomes Out and Earned Autonomy, two books on agentic AI by Dr. Harish Kotadia, Ph.D.
My books go deeper on both: Intent In, Outcomes Out and Earned Autonomy.

Go deeper

© Dr. Harish Kotadia, Ph.D., All Rights Reserved, 2026

Dr. Harish Kotadia, Ph.D., is an Enterprise AI Architect with 20+ years of IT consulting experience serving Fortune 100 clients, specializing in agentic AI systems built on Anthropic Claude, AWS Bedrock, and Google Vertex AI.

Disclaimer: This blog post is based on publicly available academic publications, vendor documentation, open standards, and news items from reputed media sources linked above. This post is intended for educational purposes, to help the enterprise agentic AI community build a shared vocabulary from public, authoritative sources.

Views and opinions expressed here are my own and do not represent those of any employer or client, past or present. The analysis presented is my independent interpretation of the published sources linked above and does not constitute legal, financial, or consulting advice of any kind.

 


Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe to get the latest posts sent to your email.

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Agentic AI Governance | Dr. Harish Kotadia, Ph.D.

Subscribe now to keep reading and get access to the full archive.

Continue reading