Skip to main content
Blog

AI Orchestration Lab: Automating the Entire Software Development Life Cycle

Last updated Agents

The appealing version of this idea is a pipeline where a requirement enters at one end and working software emerges at the other, with agents handling analysis, design, implementation, testing, review and deployment. The version that exists today is more interesting and considerably less tidy: several stages of the software life cycle automate genuinely well, several resist automation for structural reasons, and the value comes almost entirely from knowing which is which.

This guide is a survey of that landscape. It goes stage by stage through the life cycle, assesses honestly what can be delegated, explains why the resistant stages resist, and describes how to wire the automatable parts together without building a machine that produces plausible work nobody can verify.

What you will learn
  • Why verifiability, not difficulty, determines what automates well
  • A stage-by-stage assessment across the whole life cycle
  • How orchestration between agents actually works, and where it breaks
  • The human checkpoints that must remain, and why
  • What to measure so you know whether any of it helped
  • A realistic adoption sequence rather than an end-state diagram
In this article
  1. The organising principle
  2. Requirements and analysis
  3. Architecture and design
  4. Implementation
  5. Testing
  6. Code review
  7. Documentation
  8. Build and deployment
  9. Operations and incident response
  10. Maintenance and modernisation
  11. How orchestration works
  12. Where multi-agent systems break
  13. Human checkpoints
  14. Governance and security
  15. Measuring whether it helped
  16. An adoption sequence
  17. Twelve mistakes
  18. A worked example: one ticket through the lab
  19. Frequently asked questions

1. The organising principle

One property predicts whether a stage automates well, and it is not how difficult the stage is. It is whether the output can be checked cheaply and automatically.

Where a machine can verify its own work — tests pass, the build succeeds, the schema validates, the scan is clean — an agent can iterate towards correctness on its own. Where verification requires human judgement, taste, or knowledge that exists nowhere in the system, the agent produces something plausible and stops, and someone must read it carefully to discover whether it is right.

That single distinction sorts the entire life cycle:

Verifiable automaticallyRequires human judgement
Does it compile, do tests passIs this the right thing to build
Does it match the schemaIs this architecture appropriate in five years
Is the dependency vulnerableIs this trade-off acceptable to the business
Did the deployment succeedIs the user experience good
Is coverage above the thresholdDo these tests assert anything meaningful

The left column automates well today. The right column does not, and no amount of orchestration changes that — because the difficulty is not capability, it is the absence of a signal to iterate against.

2. Requirements and analysis

Automates well: summarising a long thread into a coherent statement, drafting acceptance criteria from a description, identifying missing cases by comparing against similar past work, translating between the language a customer used and the language the system uses, and generating the list of questions a requirement leaves unanswered. That last one is genuinely useful and under-used.

Resists automation: deciding what is worth building. This requires knowing the strategy, the customers, the commercial constraints, what was tried before and why it failed — most of which is not written down anywhere a system can read.

The productive pattern is the machine as a thoroughness aid rather than a decision-maker: it drafts, enumerates edge cases, and asks the questions; a person decides. The failure pattern is generating detailed specifications for work nobody validated as worth doing, which produces a well-specified waste of a quarter.

3. Architecture and design

Automates well: enumerating options with their trade-offs, drafting an architecture decision record from a discussion, checking a proposed design against known constraints, and identifying what a change would touch across a codebase. The impact analysis is particularly valuable and is the kind of tedious cross-referencing people skip.

Resists automation: the decision itself, because architecture is the accumulation of constraints that are mostly implicit. Which team maintains this. What we tried in 2022 and why it failed. What the commercial roadmap implies about scale in two years. Which vendor relationship is ending. None of that is in the repository.

The useful framing: an agent can produce a competent survey of approaches faster than a person can, and it cannot weigh them, because weighing requires knowing what your organisation actually values. Use it to ensure the option space was explored, then decide with people who hold the context.

4. Implementation

The stage where automation is most mature, and the value distribution within it is uneven.

Automates well: mechanical changes spanning many files — renaming a concept, migrating between library interfaces, adding a parameter everywhere it is used; boilerplate of every kind; translating a well-specified behaviour into code that resembles code already present; and writing the fifth similar handler when four exist as examples.

Resists automation: novel algorithmic work, anything requiring domain knowledge absent from the codebase, changes with ambiguous requirements where a person would ask a question, and work in codebases whose conventions live only in reviewers' heads.

The determining factor within implementation is again verifiability. An agent with a test suite it can run behaves entirely differently from one without: it self-corrects, converges, and reports honestly when it cannot. Without tests it produces something confident and unverified, and the confidence is unrelated to correctness.

5. Testing

A stage with a subtlety that decides whether automation helps or actively harms.

Automates well: generating tests from a specification, extending coverage into untested branches, producing data for parameterised cases, and creating characterisation tests around legacy code you are about to change. That last use is genuinely excellent and is the safest way to start.

The trap: tests generated from the implementation encode whatever the code does, including its bugs. They assert that the code does what the code does, which is tautological, and they produce coverage figures that look like assurance while providing none. This is the most damaging misapplication of automation in the entire life cycle, because it degrades the very signal everything else depends on.

The discipline: generate tests from the requirement or the specification, never from the finished implementation. Review generated tests as the specification they are — more carefully than the implementation, because the implementation is checkable and the tests are what check it.

Also automates well: property-based test generation, where you describe invariants and the machine explores inputs; and mutation testing, which measures whether your tests would actually catch a change.

6. Code review

Automates well: the mechanical layer — style, common bug patterns, missing error handling, security patterns, convention violations, and flagging changes that touch sensitive areas. Removing this layer from human review is a genuine improvement, because it lets reviewers spend attention on what only humans can assess.

Resists automation: whether the change solves the right problem, whether the approach fits the system's direction, and whether the code will be comprehensible to whoever maintains it. These require intent, which is not in the diff.

The risk specific to this stage is worth naming: as more code is generated, review becomes both more important and harder — harder because the usual signals of carelessness no longer correlate with defects. Generated code is idiomatic and confident even when wrong. A review process that does not adapt will pass things it would previously have caught.

The adaptations that work: review against the requirement rather than only the code, require the author to explain non-obvious decisions, and keep changes small. And the governing norm, stated explicitly: the person who submits a change owns it entirely, regardless of how it was produced.

7. Documentation

The stage with the best ratio of value to risk, and the most consistently neglected.

Automates well: API reference from specifications, changelogs from commit history, architecture summaries of unfamiliar modules, runbook drafts from incident history, migration guides between versions, and onboarding material describing how a system fits together. Documentation is low-risk because it is easy to verify by reading and because being wrong is embarrassing rather than dangerous.

Resists automation: explaining why something is the way it is. The most valuable documentation records decisions and their reasoning, and the reasoning exists in people's memories rather than in the code.

A pattern worth adopting: generate the description, have a person add the rationale. The generated part covers what the code does; the human part covers why, which is the part nobody can reconstruct later.

8. Build and deployment

Largely automated already, and mostly by deterministic tooling rather than by anything intelligent — which is correct. This is the stage where you least want probabilistic behaviour.

Where intelligence adds value: diagnosing why a build failed and suggesting a fix, generating a pipeline configuration for a new service from existing examples, analysing flaky test patterns across many runs, and drafting release notes from the changes included.

Where it should stay out: the deployment decision itself, and any action against production. A deterministic pipeline with defined gates is the right mechanism, and inserting judgement into it makes deployments less predictable rather than more capable.

The useful principle: automation here should explain and prepare rather than decide and act. An agent that diagnoses a failure and drafts a fix for review is valuable; one with credentials to deploy is a risk with no offsetting benefit.

9. Operations and incident response

Automates well: correlating signals across logs, metrics and traces during an incident; summarising what changed recently; drafting an initial incident timeline; searching past incidents for similar symptoms; and producing the first draft of a post-incident review from the recorded facts.

The genuine value here is speed of orientation. The first ten minutes of an incident are spent establishing what is happening, and a system that assembles the relevant signals immediately compresses that meaningfully.

Resists automation: the diagnosis itself, which requires forming and testing hypotheses about system behaviour over time, and the judgement about what to do — whether to roll back, whether the risk of a fix exceeds the risk of the current state, whether to communicate to customers.

Automated remediation deserves a specific caution. Automatically restarting a service or scaling in response to a signal is well-established and fine. Automatically applying a code change to production in response to an incident is not, because the failure mode of a wrong remediation during an incident is considerably worse than the delay it saved.

10. Maintenance and modernisation

Where the most under-appreciated value sits, because this work is genuinely tedious and genuinely verifiable.

Automates well: dependency upgrades with test verification; framework migrations following a documented pattern; language version updates; removing dead code identified by usage analysis; applying a consistent fix across many occurrences; and translating between similar technologies.

The reason this works is that success is checkable — the tests pass or they do not — and the work is repetitive enough that human attention adds little. Dependency upgrades in particular are a task most teams defer until they become a crisis, and automating them turns a quarterly ordeal into a continuous background process.

Resists automation: deciding what to modernise and in what order, which requires knowing the business trajectory; and any migration where the target requires design decisions rather than mechanical translation.

11. How orchestration works

Orchestration is what turns individual capabilities into a pipeline, and it comes in three shapes with quite different properties.

Sequential. Each stage's output feeds the next, with a checkpoint between. Simple, understandable, and error compounds through the chain — a misunderstanding at the requirements stage produces confidently wrong code three stages later.

Parallel with synthesis. Several agents attempt the same task independently, and their outputs are compared or combined. More expensive, and substantially more reliable for tasks with a wide solution space, because agreement between independent attempts is a signal.

Iterative with verification. An agent attempts, a verifier checks, and the loop continues until the check passes or a limit is reached. This is the only shape that reliably converges, and it requires an automatic verifier — which brings the discussion back to the organising principle.

The design rules that matter across all three:

  • Bound every loop — iterations, tokens, wall-clock. Unbounded loops consume real money and produce nothing.
  • Checkpoint between stages, because error compounds: a process that is ninety-five percent reliable per stage is about sixty percent reliable across ten.
  • Make every step traceable. What was decided, on what basis, by which stage. Without this an unexpected outcome is unexplainable.
  • Make everything reversible or confirmed. Anything with external consequence needs one or the other.

12. Where multi-agent systems break

Elaborate agent pipelines fail in consistent ways, and the failures are structural rather than incidental.

  • Compounding error. The arithmetic above is unforgiving and is the main reason long chains disappoint.
  • Confident handoff. Each stage passes its output as fact. A misunderstanding at stage two is treated as a requirement at stage five, and nothing in the chain expresses doubt.
  • Lost context. Summarising between stages discards precisely the detail that mattered, and the loss is invisible.
  • No verification between stages. Without a check, the pipeline optimises for producing output rather than correct output.
  • Cost invisibility. A six-stage pipeline over a hundred tickets is a substantial bill that nobody attributed to a feature.
  • Unexplainable outcomes. When the result is wrong, tracing which stage introduced the error requires instrumentation that was not designed in.

The practical conclusion: short chains with verification beat long chains without it, consistently. Two stages that each check their work outperform eight that do not, and the two-stage version is debuggable.

13. Human checkpoints

Certain decisions should remain human regardless of how capable the automation becomes, and it is worth being explicit about which and why.

CheckpointWhy it stays
Is this worth buildingRequires strategy, commercial context and history
Architectural decisions with long consequencesConstraints are implicit and mostly unwritten
Approving a change for mergeAccountability must attach to a person
Anything touching production directlyBlast radius exceeds the benefit of speed
Anything involving customer data or moneyErrors are not recoverable by retrying
Incident response decisionsJudgement under uncertainty with incomplete information
Whether tests assert anything meaningfulThe verification everything else depends on

The pattern connecting them: humans stay where the cost of being wrong is high and the signal for automatic correction is absent. That is a stable criterion rather than a temporary one.

14. Governance and security

An automated pipeline with access to your code, your infrastructure and your data needs the same treatment as any other privileged system.

  • Least privilege per stage. A documentation agent needs read access; nothing in the pipeline needs production credentials.
  • Sandboxed execution. Agents that run commands do so in isolation, with no path to production and no access to secrets.
  • No direct writes to protected branches. Everything arrives as a change proposal subject to normal review.
  • Treat retrieved content as untrusted. Instructions embedded in a file, an issue or a web page can attempt to redirect an agent's behaviour. The defence is architectural — limited tools, permissions enforced outside the model — not a prompt asking it to be careful.
  • Data policy. What may be sent to which provider, settled in writing before tools proliferate.
  • Full audit trail. What each stage did, on what input, producing what. Required for explaining an outcome and for any regulated environment.
  • Cost controls. Per-pipeline budgets that reject rather than alert.

15. Measuring whether it helped

The metrics that matter are the ones that were always the point, not new ones invented for the occasion.

MeasureWhat it tells you
Cycle time, commit to productionWhether delivery actually accelerated
Change failure rateWhether speed cost quality
Defect escape rateWhether verification kept up with volume
Review turnaround timeThe first thing to degrade when output rises
Unplanned work shareWhether quality problems are eating capacity
Cost per delivered changeWhether the automation pays for itself

Two measures to avoid actively. Lines of code produced rewards exactly the wrong behaviour. And tasks completed by agents measures activity rather than value — a pipeline that completes fifty tasks nobody needed has produced nothing.

The signal to watch most closely is review turnaround. If output rises and review capacity does not, the effective outcome is unreviewed code reaching production, which is a worse position than before any of this started.

16. An adoption sequence

  1. Settle the data and security policy. What may be sent where, what agents may access, how they are sandboxed.
  2. Start with documentation. Lowest risk, immediate value, easy to verify by reading, and it builds familiarity.
  3. Then dependency upgrades. Tedious, verifiable by tests, and it addresses work most teams defer until it is a crisis.
  4. Then test generation from specifications, with careful review of the generated tests as specifications.
  5. Then implementation of well-specified changes, in areas with good test coverage.
  6. Then review automation for the mechanical layer, freeing human attention for judgement.
  7. Then incident support — correlation and summarisation, not remediation.
  8. Only then consider chaining stages, and keep the chains short with verification between.

The sequence is deliberate: each step builds verification capability that the next depends on. Teams that start at step eight discover that they have built an elaborate pipeline producing work nobody can check.

17. Twelve mistakes

  1. Automating stages that cannot be verified. Plausible output, no way to know if it is right.
  2. Long chains without checkpoints. Compounding error, unexplainable results.
  3. Generating tests from the implementation. Encodes bugs as expected behaviour and destroys the verification signal.
  4. Increasing output without increasing review. Unreviewed code in production.
  5. Agents with production credentials. Blast radius far exceeding the benefit.
  6. Unbounded loops. Real money, no output.
  7. Treating retrieved content as trusted. Instructions can hide in a document or an issue.
  8. No cost attribution per pipeline. A large bill nobody can explain.
  9. Measuring tasks completed. Activity, not value.
  10. Automating the decision about what to build. Efficient production of the wrong thing.
  11. Removing the human from merge approval. Accountability with nowhere to attach.
  12. Starting with the most ambitious pipeline. Nothing verifiable, nothing debuggable, nothing kept.

18. A worked example: one ticket through the lab

Consider a realistic ticket — add rate limiting to a public API endpoint — moving through a pipeline that has been built with the principles above rather than as a diagram.

Analysis is assisted, not delegated. The ticket says "add rate limiting". An agent reads it alongside similar past changes and produces a list of unanswered questions: limit per what — address, account, key? What happens on exceeding it — reject, queue, degrade? Does the limit differ per plan? What should the response say? Five questions, produced in seconds, that a person then answers by talking to product and support. The agent did not decide anything; it ensured nothing was left ambiguous, which is where most rework originates.

Design is a survey, then a decision. The agent enumerates three approaches with trade-offs — at the gateway, in application middleware, or as a shared service — and identifies that one of them conflicts with an existing deployment constraint it found in the infrastructure code. The choice is made by an engineer who knows that the gateway is being replaced next quarter, which is information the agent had no access to and which decides the question.

Tests are written from the agreed behaviour, before implementation. Under the limit, at the limit, over the limit, the reset boundary, and the behaviour when the counter store is unavailable. That last case is the interesting one, and it is included because a person asked what happens if the store is down — a question the specification did not raise. These tests are reviewed carefully, because they are the specification everything downstream is verified against.

Implementation is delegated, with the tests as the verifier. The agent works in a sandbox, runs the tests, and iterates. It fails twice and reports honestly on the second failure that the reset boundary behaviour is ambiguous — which it is, because the tests permitted two readings. A person clarifies, the test is tightened, and the third attempt passes. That exchange is the loop working correctly: the agent surfaced ambiguity rather than choosing silently.

Review is human, assisted. Automated checks cover style, security patterns and dependency scanning before a person sees it. The human review finds one real issue: the implementation returns a header advising when to retry, and the value is computed from the wrong clock, so under load it advises a time in the past. Nothing about the code looks wrong. Catching it required someone thinking about behaviour rather than reading syntax.

Documentation and release notes are generated and edited. The generated text accurately describes what the limit does. An engineer adds the sentence explaining why the limit is set where it is, which is the part nobody could reconstruct in a year.

Deployment is deterministic. The pipeline runs its gates, the change goes behind a flag, and it is enabled progressively. No agent has credentials to any of this, and nothing about the process is probabilistic.

The ticket takes perhaps a day and a half rather than two and a half days. The saving is concentrated in the mechanical work, and every one of the three genuinely important moments — the deployment constraint, the store-unavailable case, and the clock bug — came from a person. That ratio is the honest picture of what this looks like when it works.

19. Frequently asked questions

Can the whole life cycle actually be automated?

Not today, and the obstacle is structural rather than a matter of capability. The stages that resist are those where output cannot be verified automatically, because that is what an agent needs in order to iterate towards correctness. Deciding what to build, weighing architectural trade-offs and judging whether tests assert anything meaningful all lack that signal.

Where should we start?

Documentation, then dependency upgrades. Both are low-risk, both are verifiable, both address work teams routinely defer, and both build the familiarity and the verification capability that later stages depend on. Starting with an end-to-end pipeline produces something impressive that nobody can check.

How many agents should a pipeline have?

Fewer than the diagrams suggest. Error compounds across stages, so two stages that each verify their output beat eight that do not — and the two-stage version can be debugged when something goes wrong. Add stages only when each one has a check it must pass.

Should agents be able to merge code?

No. Accountability must attach to a person, and merge approval is where it attaches. Agents should open change proposals subject to normal review, with no write access to protected branches. This is a governance position rather than a capability judgement, and it does not become obsolete as models improve.

What about the cost?

A multi-stage pipeline over many tickets is a substantial and easily invisible expense. Attribute cost per pipeline and per ticket in your own telemetry, set budgets that reject rather than alert, and bound every loop. Then measure cost per delivered change against what it was before — which is the only number that answers whether it paid.

Does this reduce the number of engineers needed?

It changes the composition of the work more than the quantity. The mechanical portions shrink; specification, verification, review and judgement grow, and those were always the expensive parts. Teams that reduce headcount on the assumption that production speed equals delivery speed generally discover that review became the bottleneck.

How do we stop quality degrading?

Increase review capacity before increasing output, gate on verification rather than on volume, generate tests from specifications rather than implementations, and watch defect escape rate and review turnaround as leading indicators. Quality degrades quietly in this arrangement, so the measurement has to be deliberate.

What is the single most important principle?

Automate what can be verified. Everything else in this guide follows from it — which stages to attempt, how short to keep chains, where humans must stay, and why testing generated from implementations is so damaging. If there is no automatic check, an agent cannot converge, and what you get is confident output nobody validated.

Key takeaways

  • Verifiability, not difficulty, decides what automates. Without an automatic check, an agent cannot converge.
  • Short chains with verification beat long chains without it. Error compounds unforgivingly.
  • Never generate tests from the implementation. It destroys the signal everything else depends on.
  • Documentation and dependency upgrades are the best starting points. Low risk, verifiable, genuinely tedious.
  • Humans stay where being wrong is expensive and no signal exists. A stable criterion, not a temporary one.
  • Measure delivery, not activity. Tasks completed and lines produced reward the wrong behaviour.

The useful version of an orchestration lab is not a machine that replaces the life cycle. It is a set of well-chosen automations at the stages where a machine can check its own work, connected by short chains, with people at the decisions that carry consequences — which turns out to accelerate delivery considerably more than the ambitious version does.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch