Introduction
Harrison Chase’s discussion of the agent development lifecycle (ADLC) at Interrupt26 NYC prompted me to think more carefully about where that lifecycle belongs. An agent may investigate an email, but the application decides how the investigation enters a queue, reaches an analyst, or leads to an action. Improving the investigation and changing the workflow are connected activities, with different design questions and different evidence of progress.
My view is that we should coordinate the ADLC separately but with the development of the application it powers. This distinction matters because an agent can change independently and its behavior affects both the application’s design and outcome in a tighltly coupled devlelopment loop. In this article, I explain how we can organize that loop through explicit design investigations, shared requirements, and evaluation cases that follow the agent’s results into the application. The same team may own both responsibilities, but each needs to remain visible and distinct in the development plan.
Background
Existing ADLC guidance includes substantial design and experimental work. Harrison Chase describes comparisons among prompts, models, retrieval strategies, tool schemas, and orchestration patterns within Build, Test, Deploy, and Monitor [1]. Salesforce begins with Ideation and Design and describes an inner development loop and an outer monitoring loop [2]. I build on that foundation by asking what changes when the agent is one subsystem within a larger application, with requirements and release decisions that the two development efforts must coordinate.
My earlier articles examined the scientific work behind capability development. The Missing Phase in Agentic Systems Engineering argues for explicit time and evidence before committing to a design [3], and Insight Is the Currency of Data Science examines how experimentation produces understanding of agent behavior [4]. That investigation continues as an application evolves, including occasions when evidence challenges the assumptions behind its design. Argyris’s double-loop learning similarly asks us to reconsider governing assumptions when correcting actions proves insufficient [9].
Systems engineering connects the requirements and design of individual parts to the purpose of the complete system throughout its life and is well documented and utilized. Notably, the Systems Engineering Handbook published by the National Aeronautics and Space Administration (NASA) describes this coordination across levels of a system [5]. In software, consumer-driven contracts also make a provider’s obligations visible through expectations supplied by its consumers [10]. Together these established ideas form the basis of the coordination mechanism I propose in this article.
Treating the agent as a subsystem means developing its capability explicitly and checking its contribution and impact to the application. Here we first examine the system boundary, then the agent’s development loop. The requirements and coordination sections follow and explain how shared evaluation cases connect the work, and finally the Discussion considers the costs and limits of separating the lifecycles with the Conclusion drawing out practical steps.
···
The agent is a system within the application
First things first, the basis of this article is in treating the agent as a distinct system within the application, with internal parts whose interactions shape its ability to perform a task. To better understand, Figure 1 moves from an organism to a cell to illustrate complexity can exist at more than one scale. The cell has internal processes and participates in a larger system, just as an agent has internal interactions and contributes to an application’s outcome. Drawing the agent as one box can hide a substantial development problem within it.
An agent’s harness coordinates the model, memory, tools, and any sub-agents, assembling context and controlling execution. Context supplies information for the current step, and memory retains information for later retrieval. An investigation may depend on recovering evidence gathered earlier, so performance depends on what was retained, what was retrieved, and how it reaches the model. Sub-agent handoffs introduce further questions about whether evidence survives as the work moves between parts.
Agent evaluation examines the capability produced by that composition, and application evaluation follows its results through the complete workflow. Both levels need evidence about their requirements and intended use. The application retains responsibilities for identity, authorization, and operational controls even as the agent’s internal design changes, which makes the boundary a continuing subject of development.
The agent needs a development loop
The ADLC connects development to learning from operational use, and its experimental work can include changes to the design [1, 2]. I would make the decision to reconsider that design explicit. Figure 2 shows a return from Test and evaluate to Build for changes within the current hypothesis, and a return to Plan when evidence calls the hypothesis or task decomposition into question. Planning defines the next investigation and its acceptance criteria, including any assumptions that need discussion with the application team.
An investigation that repeatedly loses evidence between sub-agents illustrates the difference. A team might improve the handoff format and evaluate the change within the current decomposition. It might also question whether dividing the investigation was useful and compare the design with a single agent. The two return paths expose that choice; teams can make either kind of change within their existing development process.
Verification and validation clarify what the evidence establishes. Verification checks conformance to specified requirements, and validation examines suitability for the intended use and environment [5]. Software engineering includes both through testing at multiple levels [7]. A unit test may verify that a tool represents missing data correctly, yet the agent may still interpret the result poorly. Agent evaluations can support verification of behavioral requirements and validation of task performance; application evaluation extends that inquiry to the complete workflow.
Readiness evidence must account for variation between runs, using repeated trials across representative cases and recording the configuration, scoring criteria, and uncertainty in the estimates. I would set acceptable error rates and the required confidence level before testing, then compare confidence bounds around the estimated rates with those limits. For consequential failures, that means asking whether the upper bound on the estimated failure rate supports the proposed scope. Repeating a few familiar cases cannot establish coverage of unfamiliar ones. Anthropic distinguishes capability evaluations that measure developing ability from regression evaluations that protect established performance [6]. An improvement needs evidence of gain alongside regression checks; maintenance may preserve capability, and an initial release needs evidence for its intended scope.
Early deployment can remain part of this process. Chase advocates controlled release and learning from use [1], which I would apply through a narrow operating scope, such as shadow mode that records proposed actions without executing them, or output that an analyst reviews before action. The application and agent teams can widen that scope as evidence accumulates, with monitoring returning failures and new cases to development.

Requirements connect the agent to the application
The application’s intended outcome determines what the agent needs to accomplish and how its results will be used. Consider an email investigation that reaches an inconclusive result because evidence is unavailable. If the application forces every result into a safe or malicious label, it can turn an appropriate expression of uncertainty into an unsupported decision. The agent needs to preserve what it established and what remains unknown, and the application needs a suitable next step.
A versioned behavioral contract can express those shared expectations as integrated evaluation cases. I would make the application team accountable for the acceptance criteria and have both teams maintain the cases. Each case records the task and available evidence, permitted agent outcomes, expected application action, and scoring rules. For a case with unavailable reputation data and no other evidence that resolves it, the agent should return an inconclusive result and the application should send it for review without automatically releasing the message. Repeated trials follow the case through the complete workflow, including failures of the review path.
Review capacity makes the requirement quantitative. If an illustrative application processes 10,000 messages a day and has 200 review slots available, a 2% inconclusive rate would consume all of them, leaving no headroom for bursts or other referrals. The teams need a lower operating target and a response to excess demand, such as narrowing automated scope or increasing capacity. Reducing the queue by forcing confident classifications would defeat the requirement.
Reliability measures also belong in the contract because the application’s use determines what success means. Anthropic describes pass@k as the chance of at least one success in k attempts and passk as the chance that all k attempts succeed [6]. More attempts can improve the first measure and reduce the second. For automated email decisions, I would specify per-case consistency and false-safe error bounds under the actual retry policy, since occasional success across several attempts cannot justify acting on every result.
Coordinate the two development lifecycles
The behavioral contract connects independent development work to a shared release decision. Following the consumer-driven contract principle [10], the application team supplies the expectations its workflow depends on, and the agent team runs those cases against candidate changes. The contract adds statistical acceptance criteria and integrated outcomes to interface checks. Both teams review changes to the contract itself, so a failing candidate prompts investigation or an explicit requirements decision.
Behavior-changing updates need this check regardless of how they are delivered. Chase’s context hub example allows prompts and context to change without a full deployment [1]. I would therefore run the agreed cases before promoting changes to prompts, context configuration, models, tools, or code, recording their versions with the results. Integrated checks should begin with a minimal working path through the application and expand with its scope. A shared release gate then considers the intended operating scope and the evidence from both teams, as Figure 3 shows.
The system architect needs defined decision rights to keep that coordination workable. I would assign the architect responsibility for requirement allocation, interface meaning, and review of changes that affect both sides, with the application owner accountable for operational acceptance. Teams can make changes within those agreements independently. Figure 3 represents application work through a simplified software development lifecycle (SDLC), coupled to the agent cycle through the behavioral contract and integrated evaluation.

A system structure view locates these responsibilities in the application. Figure 4 uses nested parts inspired by Systems Modeling Language (SysML) v2 [8], showing the harness, model, memory, tools, and optional sub-agents within the agent. Identity and operational controls span the application, with authorization enforced at tool access. The boxes describe logical responsibilities that can guide development even when components share infrastructure.

···
Discussion
A separate agent lifecycle where uncertainty about capability needs sustained investigation of where agent changes have consequences that application delivery can obscure. Figure 1 helps us keep both the agent’s internal design and its role in the application in view. For a low-consequence feature with a single model call and tightly constrained handling, one team may manage the evaluation within its ordinary workflow. A single call can still justify substantial evaluation when its output controls an important decision, so the choice depends on the consequences and uncertainty involved.
Coordination has costs that we need to account for. Repeated trials consume time and compute, shared cases require maintenance, and an architect who approves every change can become a bottleneck. I would start with cases that exercise the important interactions, automate routine checks, and reserve joint decisions for changes to requirements, interface meaning, or operating scope. A small team can carry both responsibilities, and larger teams can divide them as the work warrants.
The lasting benefit, for me, is that our understanding can accumulate across changes in technology. A replacement model may require a new investigation even when the application backlog is unchanged, but we can begin with the requirements, failure cases, and design observations we have preserved. This is where the scientific work in my earlier articles [3, 4] connects to everyday delivery. We can adopt new capabilities and retain what experience has taught us about the problem we are trying to solve.
Conclusion
In this article, we examined why agent development should be managed as a distinct lifecycle within the development of the application it powers, with its own requirements, ownership, and evidence of capability. Evaluation should inform both how we improve a design and when we reconsider its underlying assumptions, making the return to planning an explicit part of that lifecycle. The agent and application may be developed by different teams and progress at different rates, but their development must remain closely coordinated through shared requirements, interfaces, and evaluation. Systems engineering provides a foundation for that coordination, helping us preserve distinct development efforts and hold them accountable to the performance of the whole system.
References
[1] Chase, H. (2026, May 9). The agent development lifecycle (ADLC). LangChain. https://www.langchain.com/blog/the-agent-development-lifecycle
[2] Salesforce Architects. (n.d.). The agent development lifecycle: From conception to production. https://architect.salesforce.com/docs/architect/fundamentals/guide/agent-development-lifecycle
[3] Hinton, A. (2026, September 5). The missing phase in agentic systems engineering. LinkedIn. https://www.linkedin.com/pulse/missing-phase-agentic-systems-engineering-andrew-hinton-phd-fyqse
[4] Hinton, A. (2026, September 30). Insight is still the currency of data science. Towards Data Science. https://towardsdatascience.com/insight-is-still-the-currency-of-data-science/
[5] National Aeronautics and Space Administration. (2016). NASA systems engineering handbook (NASA/SP-2016-6105 Rev 2). https://www.nasa.gov/reference/2-0-fundamentals-of-systems-engineering/
[6] Anthropic. (2026, January 9). Demystifying evals for AI agents. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
[7] IEEE Computer Society. (n.d.). Chapter 4: Software testing. SWEBOK Guide. Retrieved September 27, 2026, from https://swebokwiki.org/Chapter_4:_Software_Testing
[8] Object Management Group. (2025). Systems Modeling Language: Version 2.0, part 1, language specification. https://www.omg.org/spec/SysML/2.0/Language/PDF
[9] Argyris, C. (1977). Double loop learning in organizations. Harvard Business Review, 55(5), 115–125. https://hbr.org/1977/09/double-loop-learning-in-organizations
[10] Robinson, I. (2006, June 12). Consumer-driven contracts: A service evolution pattern. MartinFowler.com. https://martinfowler.com/articles/consumerDrivenContracts.html

