Predictions about artificial intelligence and software development have been consistently wrong in both directions. The claim that programming would be obsolete within two years was wrong. So was the claim that these tools were autocomplete with better marketing. What has actually happened is stranger and more interesting: the cost of producing code has fallen substantially, while the cost of deciding what code should exist, and verifying that it is correct, has not moved at all.
This guide examines what has genuinely changed in how software is built, what the evidence says about productivity, which parts of the job are being reshaped, and what the realistic horizon looks like. It avoids both the enthusiasm and the dismissal, because neither is a useful basis for planning.
What you will learn- What AI tooling actually does well in software development today
- Why the productivity picture is more complicated than the headlines
- How coding agents work, and where they reliably break down
- What changes about code review, testing and architecture
- The new failure modes teams are encountering, and how to manage them
- How roles are shifting, and what skills become more valuable
- What has actually changed
- The four capability tiers
- The productivity question, honestly
- How coding agents work
- Where they break down
- What changes about code review
- Testing in an AI-assisted workflow
- Architecture and the shifting cost curve
- New failure modes
- Security implications
- What this does to roles
- Adopting it without damage
- What is plausible next
- Twelve mistakes
- A worked example: one feature, assisted
- Frequently asked questions
1. What has actually changed
Three things, precisely.
Writing code became cheaper. Not free, and not automatic, but substantially cheaper for the large category of work that is well-specified and resembles code that already exists. Boilerplate, tests, data transformations, API clients, configuration, migrations between similar shapes — this is a meaningful share of what most engineers spend their days on.
Reading unfamiliar code became easier. An underrated shift. Explaining what a piece of code does, summarising an unfamiliar module, tracing why something behaves as it does — these were previously slow and are now fast. For anyone joining a codebase or maintaining a system they did not build, this may be the larger benefit.
The cost of a first attempt collapsed. Producing a rough implementation to react to is now nearly free. That changes how exploration works: rather than reasoning about which approach might be right, it is often faster to generate three and look at them.
What has not changed is equally specific. Deciding what to build, understanding a business domain, choosing between architectures with long-lived consequences, debugging a subtle production incident, and verifying that a system is actually correct all remain roughly as difficult as they were. And since these were already the expensive parts of software engineering, the overall effect is smaller than the change in code production alone would suggest.
The clarifying frame: these tools have made the typing faster. Software engineering was never primarily typing.
2. The four capability tiers
| Tier | What it does | Reliability |
|---|---|---|
| Completion | Suggests the next lines as you type | High for routine code; low value for novel logic |
| Conversation | Answers questions, explains code, drafts functions on request | Good, with verification; confident when wrong |
| In-editor agents | Multi-file changes across a codebase with your supervision | Strong on well-specified tasks; needs review |
| Autonomous agents | Take an issue, plan, implement, test, open a change for review | Works on bounded tasks; degrades over long sequences |
The value of each tier depends heavily on the task. Completion is most useful in verbose languages and repetitive code. Conversation is most useful for unfamiliar territory. In-editor agents shine on mechanical refactoring across many files — the kind of change that is conceptually trivial and tediously large. Autonomous agents work best where success is verifiable by tests, because verification is what bounds the error.
3. The productivity question, honestly
Reported gains vary enormously across studies and organisations, and the variation itself is informative. The pattern that emerges is consistent even when the numbers are not.
Gains are largest for tasks that are well-specified, resemble abundant existing code, are verifiable by tests, and are performed by someone unfamiliar with that particular technology. Writing a standard integration in a language you rarely use is where the improvement is dramatic.
Gains are smallest, and occasionally negative, for tasks in large unfamiliar codebases with implicit conventions, work requiring deep domain knowledge, novel algorithmic problems, and debugging subtle behaviour. Experienced engineers working in a codebase they know well sometimes find that reviewing generated code takes longer than writing it themselves.
Two systematic biases distort the reported picture. Perceived speedup consistently exceeds measured speedup — people feel faster because the visible activity of typing decreased, even when total time did not. And measurement almost always stops at code produced rather than at working software delivered, which omits review, integration, defect and maintenance costs.
The honest summary: a real and useful improvement in a substantial category of work, considerably smaller than marketing suggests, and highly dependent on task and context. Organisations reporting transformational gains are usually measuring code volume; those reporting none are usually applying the tools to work they do not suit.
4. How coding agents work
Understanding the mechanism explains the failure modes. An agent operates in a loop: it receives a goal, decides what information it needs, requests a tool, receives the result, and repeats until it believes the task is done.
The tools are ordinary: read a file, search the codebase, edit a file, run a command, execute tests. The model never touches your systems directly — it asks, and your code decides whether and how to comply. Everything an agent can do is something you explicitly gave it the ability to do.
Three factors determine whether the loop converges on something useful.
Context quality. The agent can only work with what it can see. A codebase with clear structure, meaningful names, a good readme and existing tests gives it far more to work from than one where conventions live in people's heads.
Verification availability. An agent that can run tests knows whether it succeeded. One that cannot is guessing, and its confidence is unrelated to its correctness. This single factor explains most of the variance in agent usefulness across projects.
Task boundedness. Error compounds across steps. A process that is ninety-five percent reliable per step is roughly sixty percent reliable across ten. This is why bounded tasks with clear completion criteria succeed and open-ended ones drift.
5. Where they break down
- Implicit conventions. Every mature codebase has rules nobody wrote down — this is how we handle errors, this is where validation belongs, we never call that service directly. Generated code violates them plausibly, and reviewers who are skimming let it through.
- Cross-cutting changes. A change touching thirty files with subtle interactions between them exceeds what an agent reliably tracks, and partial completion is worse than no attempt.
- Ambiguous requirements. A human engineer asks a question. An agent picks an interpretation and proceeds confidently, and the divergence is discovered much later.
- Deep debugging. Diagnosing a race condition or a memory leak requires forming and testing hypotheses about system behaviour over time. Agents can help gather evidence and are poor at the reasoning that connects it.
- Performance work. Optimisation requires measurement, and agents tend to apply plausible optimisations without evidence that the bottleneck is where they assumed.
- Anything with no verification. Without tests or a way to check, output quality is unknowable and the loop cannot self-correct.
6. What changes about code review
Review becomes both more important and harder, which is an uncomfortable combination.
More important because more code arrives, generated from patterns rather than from understanding, and because the author may not fully understand it either. Review is now the primary place where correctness is established.
Harder because a reviewer's usual heuristics degrade. Generated code looks confident and idiomatic even when wrong, so the surface signals that used to indicate carelessness — inconsistent naming, awkward structure, missing error handling — no longer correlate with defects the way they did.
Three adaptations help. Review the specification alongside the code: the most common defect is code that correctly implements the wrong thing, which cannot be detected by reading the code alone. Require the author to explain non-obvious decisions, which surfaces cases where nobody understands the change. And keep changes small, which mattered before and matters more now that generating a thousand lines is as easy as generating fifty.
The failure pattern to watch for is review becoming a rubber stamp because the volume is unmanageable. If review throughput has not increased alongside code production, the effective outcome is unreviewed code reaching production — which is a considerably worse position than before.
7. Testing in an AI-assisted workflow
Tests become more valuable, for a straightforward reason: they are the verification that makes generation trustworthy. A codebase with a good suite can absorb generated changes safely. One without has no way to distinguish working code from plausible code.
Generating tests is one of the strongest applications of the technology, with one important caveat. Tests generated from the implementation encode whatever the implementation does, including its bugs — they assert that the code does what it does, which is tautological. Tests generated from the specification, or written before the implementation, retain their value as an independent check.
Two practices that become more valuable: property-based testing, which describes invariants rather than examples and catches cases nobody enumerated; and contract testing at service boundaries, which catches the integration mistakes generated code makes when it assumes an interface shape.
8. Architecture and the shifting cost curve
When the cost of writing code falls but the cost of understanding it does not, the economics of architecture shift in specific ways.
Duplication becomes cheaper to create and no cheaper to maintain. The traditional argument against duplication was partly the cost of writing it. That argument weakens; the argument about divergence and inconsistency does not. Expect more duplication unless it is actively resisted.
Explicit beats clever. Code that is verbose and obvious is easier for both humans and models to work with. Dense, idiomatic code that relies on implicit knowledge is where generated changes go wrong.
Conventions must be written down. A codebase where the rules exist only in reviewers' heads cannot communicate them to a tool. Documented conventions, linting rules that encode them, and clear examples all become more valuable — they are now instructions to a collaborator as well as guidance to a colleague.
Modularity pays more. An agent working within a well-bounded module with a clear interface has a tractable problem. One working across a tangled codebase does not. The traditional benefits of good boundaries now include machine comprehensibility.
9. New failure modes
- Plausible-but-wrong code. Correct-looking implementations with subtly wrong behaviour, which pass casual review precisely because they look right.
- Convention drift. Each generated change is individually reasonable and collectively inconsistent, and consistency erodes without any single decision to abandon it.
- Unowned code. Nobody in the team fully understands a module, because it was generated and skim-reviewed. This is discovered during an incident.
- Dependency inflation. Models suggest well-known libraries readily, and codebases accumulate dependencies for functions that were twenty lines.
- Test theatre. High coverage from generated tests that assert current behaviour and would not catch a regression that matters.
- Review fatigue. Volume rises, attention per change falls, defects pass through.
- Skill atrophy. Engineers who never debug from first principles gradually lose the ability, and it is exactly the skill an incident demands.
- Hallucinated interfaces. Calls to functions and parameters that do not exist, caught by compilation in typed languages and by users in dynamic ones.
10. Security implications
Three distinct concerns, frequently conflated.
Generated code quality. Models produce code resembling their training data, which includes a great deal of insecure code from public sources. Common issues include missing input validation, unsafe query construction and permissive defaults. The mitigation is the security tooling you should already have: static analysis, dependency scanning and secret detection in the pipeline, applied to all code regardless of origin.
What you send. Code sent to an external service is a data transfer subject to your normal classification rules. Establish whether inputs may be retained or used for training, whether processing location matters for your obligations, and what your contract actually says. This is a procurement question with a clear answer, not an unknowable risk.
Agent permissions. An agent that can execute commands has the permissions you gave it. Treat it as an untrusted process: no production credentials, no ability to push directly to protected branches, no access to secrets, and a sandbox for execution. Also treat retrieved content as potentially adversarial — instructions embedded in a file an agent reads can attempt to redirect its behaviour.
11. What this does to roles
The most likely trajectory is a shift in the composition of engineering work rather than a reduction in the need for engineers.
More valuable: specifying problems precisely, reviewing critically, system design, debugging from first principles, understanding the business domain, and building the verification that makes generation safe. Every one of these is about judgement rather than production.
Less valuable: memorised syntax, mechanical translation between formats, boilerplate composition, and knowing a specific library's interface by heart. These were always the least interesting parts of the job.
The genuine concern is the entry route. Junior engineers historically built judgement by doing the routine work these tools now absorb, and it is not obvious what replaces that apprenticeship. Organisations that use these tools to eliminate junior roles entirely may find in five years that they have no senior engineers either. The teams handling this well are shifting junior work towards reviewing, testing and debugging with strong mentorship, which builds judgement faster than boilerplate ever did.
12. Adopting it without damage
- Decide the data position first. What may be sent where, under what agreement. Answer this before tools proliferate, because retrospective policy is much harder.
- Start where verification exists. Areas with good tests, so quality is measurable rather than assumed.
- Write your conventions down. They are now instructions to a tool, not just guidance for reviewers.
- Strengthen review before increasing volume. Otherwise you have automated the production of unreviewed code.
- Measure delivered outcomes, not code produced. Cycle time, change failure rate, defect escape rate. Code volume is not an achievement.
- Set the accountability rule explicitly. The engineer who submits a change owns it entirely, regardless of how it was produced. This must be stated, because it is the norm that prevents most of the failure modes above.
- Protect the learning path. Deliberate mentorship and problem-solving practice for less experienced engineers.
13. What is plausible next
Reasonably predictable. Cost per unit of capability continues falling, which matters more commercially than capability increases. Agents become better at bounded, verifiable tasks — dependency upgrades, migrations, test generation, well-specified bug fixes. Tooling for evaluating and monitoring generated output matures. Development environments assume agent participation rather than bolting it on.
Genuinely uncertain. Whether reliable long-horizon autonomy is achievable without fundamentally new approaches. Whether the reliability curve across many steps improves enough to change what is delegable. How the profession solves the apprenticeship problem. What happens to code quality across the industry over a decade of assisted production.
Worth ignoring. Confident timelines for the end of programming, in either direction. The historical record for such predictions is poor, and the people closest to the work are consistently the most cautious.
The stable planning assumption: the constraint on software delivery moves further towards knowing what to build and verifying that it works. Organisations that invest in specification quality, testing and review will extract far more from these tools than those that simply generate more code.
14. Twelve mistakes
- Measuring lines of code as productivity. Rewards exactly the wrong behaviour.
- Increasing output without increasing review capacity. Unreviewed code in production.
- Generating tests from the implementation. Encodes existing bugs as expected behaviour.
- Accepting code nobody understands. Discovered during an incident.
- No written conventions. Nothing to guide generation, so it drifts.
- Agents with production credentials. Treating an untrusted process as trusted.
- Skipping security scanning on generated code. It needs it more, not less.
- Using it in a codebase with no tests. No way to distinguish working from plausible.
- Eliminating junior roles. Solving this year's cost at the expense of the next decade's capability.
- Long autonomous runs without checkpoints. Compounding error with no correction.
- No data policy. Confidential code sent to services whose terms nobody read.
- Believing the perceived speedup. Measure delivery, not the feeling of velocity.
15. A worked example: one feature, assisted
Consider a realistic task: adding a scheduled export capability to an existing application. It is well-specified, touches several layers, and has clear acceptance criteria. Here is where assistance genuinely helps and where it does not.
Specification is unchanged. Deciding what the feature should do — which report types, what happens when generation fails, whether recipients must be verified addresses, how permissions apply — is a conversation with product and support. No tool shortens this, and getting it wrong makes everything downstream worthless.
Exploration is dramatically faster. Asking for three approaches to scheduling — a cron-style poller, a queue with delayed messages, a dedicated scheduler service — with the trade-offs of each takes two minutes and produces a useful basis for a decision that would previously have required reading documentation for an hour. The decision itself remains a human one, informed by operational context the model does not have.
The boilerplate disappears. The migration, the entity, the repository, the API endpoints, the validation and the serialisation are generated in minutes rather than hours. This is genuinely the majority of the typing and a minority of the thinking. Review is quick because the code is conventional and the tests confirm it.
The tests are written first, by hand or from the specification. Both acceptance scenarios — the happy path and the generator-failure path — are expressed before implementation. This matters more than usual: they are what allows the generated implementation to be trusted, and generating them from the finished code instead would have produced tests asserting whatever the code happened to do.
The subtle part stays human. Timezone handling for a schedule set at seven in the morning in a user's local zone, across a daylight saving transition, is exactly the kind of problem where a plausible implementation is wrong in a way nobody notices for six months. It requires someone to think carefully about what the user means, and to write a test that pins the intended behaviour.
The failure path is where review earns its keep. The generated implementation handles the success case cleanly and, on first attempt, retries a failed generation without limit — a plausible pattern that would eventually send a customer forty identical failure notifications. Nothing about the code looks wrong; it looks careful. Catching it required a reviewer thinking about behaviour rather than reading syntax.
The net effect on this feature is perhaps a day saved on a four-day task, concentrated entirely in the mechanical portions, with the difficult parts unchanged. That ratio is typical, and it is worth having — it is simply not the transformation the marketing describes.
16. Frequently asked questions
Will AI replace software engineers?
It is changing what the job consists of rather than removing it. Producing code is a smaller share of the work than people outside the field assume; specifying, verifying, designing and debugging are larger, and these have not become easier. What is genuinely at risk is work that was purely mechanical translation. What is growing is work requiring judgement about what should exist and whether it is correct.
Should we let AI write production code?
You almost certainly already do, and the question is how it is governed. The workable position is that generated code is subject to exactly the same standards as any other: reviewed, tested, scanned, and owned by the engineer who submits it. What does not work is a separate, lower standard for generated code, or a policy of prohibition that is quietly ignored.
How do we stop convention drift?
Encode conventions mechanically wherever possible — linting rules, formatters, architectural tests that fail when a boundary is violated. Written conventions are better than unwritten ones, and enforced conventions are better than written ones. This was always true; the volume of generated code makes it urgent rather than aspirational.
What about licensing of generated code?
The legal position is unsettled and varies by jurisdiction. Practical risk management: use tools offering indemnification where the exposure matters, enable any filtering that suppresses close reproduction of training data, and avoid generating substantial code in domains where you know licensed implementations dominate. For most business software the practical risk is low; for anything you intend to license or open-source, take advice.
How should junior engineers use these tools?
Deliberately, with an emphasis on understanding rather than output. A useful discipline is to attempt a problem first, then compare with a generated solution and investigate the differences — which teaches far more than accepting the first suggestion. What builds judgement is the explanation, the debugging and the review, not the production. Teams that pair this with genuine mentorship report the transition works; those that hand juniors a tool and a ticket queue do not.
Do these tools work on large legacy codebases?
Less well, and for understandable reasons: implicit conventions, sparse tests, and structure that exceeds what any context window holds. They remain genuinely useful for explanation — understanding what an unfamiliar module does is one of the strongest applications. For modification, results improve substantially with the same investments that help humans: characterisation tests, documented boundaries and incremental modularisation.
What should we measure to know if it is working?
Delivery outcomes, not activity. Cycle time from commit to production, change failure rate, defect escape rate, and time spent on unplanned work. If code volume rises while these are flat or worse, the tools are producing work rather than value. Also worth tracking: review turnaround time, which is the first thing to degrade when volume increases without capacity.
What is the single most important adaptation?
Investing in verification. Tests, contracts, static analysis and review capacity are what convert cheap code production into cheap working software. Without them, generation increases the rate at which plausible code enters your system, and plausible is not the same as correct — a distinction your users will make for you if you do not make it first.
A team adoption checklist
These questions establish whether a team is positioned to benefit or to accumulate debt. Each is answerable with evidence rather than opinion, and the pattern of answers is more informative than any individual one.
| Question | Healthy answer |
|---|---|
| Is there a written data policy for what may be sent to which tools? | Yes, agreed before tools proliferated |
| Can the codebase verify itself? | Tests that genuinely fail when behaviour breaks |
| Are conventions written down and mechanically enforced? | Linting and architectural rules, not folklore |
| Has review capacity increased alongside output? | Yes, or output was deliberately not increased |
| Who owns a generated change? | The engineer who submitted it, without qualification |
| Do agents have production credentials? | No — sandboxed, no secrets, no protected branches |
| Does security scanning apply to all code equally? | Yes, regardless of how it was produced |
| Are you measuring delivery or activity? | Cycle time and failure rate, not lines produced |
| Is there a deliberate learning path for juniors? | Mentored review and debugging, not a ticket queue |
| Would anyone notice if a module was understood by nobody? | Yes — review requires an explanation |
The two rows that predict trouble most reliably are review capacity and verification. A team that increases output without either is not becoming faster; it is moving the cost from writing to debugging, where it is considerably higher and arrives later.
Key takeaways
- Production got cheaper; specification and verification did not. That asymmetry explains almost everything.
- Gains depend on the task. Largest for well-specified, verifiable, conventional work; smallest for novel or deeply contextual work.
- Verification bounds usefulness. Agents with tests self-correct; agents without them guess confidently.
- Review becomes the bottleneck. Increase capacity before increasing output, or you have automated unreviewed code.
- Write conventions down. They are now instructions to a collaborator, not folklore for reviewers.
- Protect the apprenticeship. Eliminating junior work eliminates the pipeline that produces senior judgement.
The most useful stance is neither adoption for its own sake nor resistance. It is to identify precisely where in your work the answer is cheap to check, apply the tools aggressively there, and invest the time you save in the parts that were always the hard bit — deciding what to build, and knowing whether it actually works.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch