An AI coding agent is given a task in a repository and, some minutes later, presents a set of changes across several files with the test suite passing. It is a genuinely different thing from a code-completion tool, and the difference is not that the model is smarter — it is that the model is inside a loop that lets it observe consequences.
This article explains that loop in detail: how an agent forms a plan, how it explores an unfamiliar codebase, how it decides what to edit, how it uses tests as a feedback signal, and where the whole arrangement breaks down. Understanding the mechanism is what turns these tools from unpredictable to reliable.
What you will learn
- The agent loop, and why observation is the key ingredient
- How agents explore a codebase they have never seen
- What planning actually consists of, and when it helps
- Why tests are the single largest determinant of quality
- The characteristic failure modes and their causes
- How to write tasks and set up repositories so agents succeed
- What makes an agent different
- The loop
- The tools an agent has
- Exploration: reading an unfamiliar codebase
- Context, and its limits
- Planning
- Making edits
- Tests as the feedback signal
- Iterating on failure
- Knowing when to stop
- The characteristic failure modes
- Permissions and blast radius
- Writing a task an agent can execute
- Preparing a repository for agents
- Reviewing agent output
- Where agents genuinely help, and where they do not
- Twelve mistakes
- A worked example: a bug fix, step by step
- Frequently asked questions
1. What makes an agent different
A completion tool predicts the next lines given what you have written. A chat assistant answers a question given the code you showed it. Both produce text and stop.
An agent does something categorically different: it acts, observes the result, and decides what to do next. It runs a search and reads what came back. It edits a file and runs the tests. It reads the failure and revises the edit.
That feedback is the whole difference. A model producing code with no ability to run it is guessing, however good the guess. A model that can run the tests is checking. The underlying language model is the same; the loop around it converts an educated guess into something verified.
This also explains the sharpest predictor of agent usefulness: agents are as good as your ability to verify their work. In a repository with a fast, trustworthy test suite, an agent iterates to correctness. In a repository without one, it produces plausible changes and cannot tell whether they work — and neither can you, faster than reading everything.
2. The loop
Every coding agent, whatever its interface, runs approximately this cycle:
- Read the task and whatever context is supplied.
- Decide on an action — search, read a file, edit, run a command.
- The harness executes it. Not the model; the surrounding program.
- Observe the result — file contents, search hits, test output, an error.
- Decide again with that result now in context.
- Repeat until the task appears complete or a limit is reached.
Step three deserves emphasis because it is the security model in one line. The model never touches your filesystem, your network or your shell. It emits a structured request; the harness decides whether to honour it and performs the action. Every permission question belongs there.
Everything the agent has observed accumulates in its context. This is why a long agent run degrades: by the fortieth action, the context holds thirty-nine previous results, and the earliest ones — which may include the task description's nuances — compete with a large volume of file contents and test output.
3. The tools an agent has
| Tool | What it does | Why it matters |
|---|---|---|
| Search | Find text or patterns across the repository | The primary means of orientation in unfamiliar code |
| Read file | Retrieve a file or a portion of one | Detailed understanding of specific code |
| List directory | See structure | Understanding project layout and conventions |
| Edit file | Apply a change | The actual work |
| Run command | Execute tests, builds, linters | The verification signal — the most important tool |
| Version control | Diff, status, history | Seeing its own changes; understanding recent context |
The set is deliberately small, and its power comes from composition rather than from any individual tool. Search then read then edit then test is the pattern of almost every task, and an agent lacking any one of those four is substantially less capable.
The tool that most changes outcomes is run command. An agent that cannot execute anything is a sophisticated text generator with repository access. An agent that can run the tests is participating in the same loop a developer does.
4. Exploration: reading an unfamiliar codebase
The first thing a competent agent does with a non-trivial task is orient itself, and how it does this is worth understanding because it tells you what to optimise for.
The typical sequence: search for terms from the task description; find a handful of candidate files; read the most promising; follow references outward — what calls this, what does it call; look for a similar existing feature to imitate; and check the test files, which frequently explain intent better than the implementation.
Three properties of a repository make this dramatically easier or harder.
Naming. An agent finds code by searching for words. A codebase where a concept is called the same thing everywhere is navigable; one where the same concept is a "customer" in one module, a "client" in another and an "account" in a third requires the agent to discover that mapping, and it frequently does not.
Locality. Related code near each other means one directory listing orients the agent. Logic scattered across a deep hierarchy means many reads, each consuming context.
An existing example. The single most useful thing a repository can contain is a feature similar to the one being asked for. Agents imitate well, and "do this the way X does it" converts an open-ended design problem into a pattern-matching one.
5. Context, and its limits
Everything the agent reads occupies context, and context is finite even when very large.
This produces a real tension. Reading more files means better understanding and less room for the reasoning and the accumulated tool results. An agent that reads twenty files thoroughly may run out of room before finishing the edits.
Competent agents manage this by reading selectively — searching to narrow before reading, reading portions of files rather than whole ones, and summarising what they learned rather than retaining raw contents. But the constraint is real, and it explains a behaviour that otherwise looks like carelessness: an agent that seems to have forgotten something you told it at the start of a long run has, in a practical sense, done exactly that.
The implication for how you work with them: smaller tasks produce better results, not because the agent cannot handle a big one, but because a big one exhausts the room available for careful work. Two focused tasks routinely beat one broad one.
6. Planning
For anything beyond a small edit, an agent benefits from producing a plan before acting — and this is not merely organisational. Generating a plan is reasoning done before committing to a first edit, and it catches misunderstandings while they are still cheap.
A useful plan identifies which files need changing and why, what order the changes must happen in, what could break, and how the result will be verified.
Where planning genuinely helps: multi-file changes, anything touching a data model, refactoring, and tasks where the requirements have interacting constraints. Where it adds little: single-file edits, mechanical changes, and anything where the work is obvious once the right file is found.
The underrated benefit of a plan is that you can read it. A plan that misunderstands the task is visible in thirty seconds, before the agent has produced a four-hundred-line diff based on that misunderstanding. For any task where you are not certain the description was unambiguous, asking for the plan first is the cheapest quality control available.
7. Making edits
Modern agents edit by targeted replacement rather than by rewriting whole files, and this matters for reasons beyond efficiency.
A targeted edit — replace this exact text with that text — is verifiable. If the original text is not found, the edit fails loudly rather than producing something subtly wrong. Rewriting a whole file means regenerating code the agent was not asked to change, which is where unrequested modifications creep in.
Two behaviours worth knowing about. Agents will sometimes edit and immediately re-read to confirm the result, which is a reasonable check and consumes context. And agents working across many files will sometimes make an inconsistent set of changes — correct individually, mismatched collectively — because each edit was decided with different portions of the codebase in view. This is the most common source of an agent's diff that passes tests and is nonetheless wrong.
8. Tests as the feedback signal
This is the section that matters most, because test quality determines agent output quality more than any other factor including the model.
An agent with tests operates in a closed loop: make a change, run the tests, read the result, and either proceed or revise. Failure output is enormously informative — it names the assertion, the expected and actual values, and often the line. That is a precise signal the agent can act on.
An agent without tests operates open-loop. It makes a change that looks right, has no way to confirm it, and reports completion based on its own judgement. Sometimes correct, entirely unverified.
Three properties of a test suite determine how well this works:
Speed. An agent may run the suite ten or twenty times in a single task. A two-minute suite makes that practical; a twenty-minute one makes it impossible, and the agent will run it once at the end rather than iterating.
Reliability. A flaky suite is worse than no suite for an agent, because it will chase a failure that was not caused by its change, sometimes modifying working code to accommodate a random failure.
Specificity. A failure saying which assertion broke and why lets the agent fix it. A failure saying only that something went wrong sends it searching.
9. Iterating on failure
What an agent does with a failing test is where competence shows.
The good pattern: read the failure, form a hypothesis about the cause, check that hypothesis by reading the relevant code, make a targeted fix, re-run. Each cycle should narrow the problem.
Two bad patterns are common enough to watch for.
Thrashing. The agent makes a change, the test still fails, it reverts and tries something else, and it cycles without converging. This almost always indicates the agent has misunderstood something fundamental — the wrong file, the wrong layer, a misread requirement — and no amount of further iteration will fix it. Thrashing is a signal to stop and re-specify, not to wait longer.
Weakening the test. The agent modifies the test so it passes. This is not deception; the objective was stated as "make the tests pass" and it did. It is nonetheless the most dangerous failure mode, because it produces a green suite that verifies nothing. Always check whether test files changed in an agent's diff, and treat any change to an assertion as requiring justification.
10. Knowing when to stop
An agent has to decide the task is done, and this judgement is genuinely difficult.
The signals available are weak: tests pass, the requested change appears made, no obvious errors remain. None of them establishes that the task was understood correctly, that the edge cases are handled, or that the approach fits the codebase.
Practical harnesses impose limits — a maximum number of actions, a time bound, a token budget — which converts an infinite loop into a truncated attempt. That is the right trade, and it means an agent that hits a limit has not necessarily failed; it may have been most of the way through something worth continuing.
The corollary for users: an agent reporting success is reporting that its checks passed, which is a narrower claim than it sounds. Review accordingly.
11. The characteristic failure modes
Solving the wrong problem. The task description was ambiguous, the agent chose an interpretation, and it executed that interpretation well. This is the most common failure and the easiest to prevent, because a plan reviewed before execution catches it.
Local correctness, global inconsistency. Each file's change is sensible; together they introduce a second way of doing something. Caused by making decisions with partial context.
Ignoring an unwritten convention. The codebase has a rule nobody documented — this layer never calls that one, this field is always set together with that one. The agent cannot see it and violates it.
Over-building. Configuration options, abstraction layers and error handling for cases that will never occur. The agent has no cost pressure and no sense of what is proportionate.
Adding a dependency for something the codebase already handles, because the agent did not find the existing utility.
Optimism about the unhappy path. Generated code frequently assumes things succeed. The failure branches are where the gaps are.
Getting stuck on environment problems. A test failing for reasons unrelated to the code — a missing service, a stale cache, a configuration issue — sends the agent into a long unproductive investigation. It cannot easily distinguish "my change is wrong" from "this environment is broken".
12. Permissions and blast radius
An agent that can run commands can run destructive ones, and the defence is environmental rather than instructional.
Version control is the primary safety net. An agent working in a clean repository can do very little that cannot be undone by discarding changes. This single property removes most of the risk, and it is why "commit before starting an agent" is worth making a habit.
Approval for irreversible actions. Deleting files, force-pushing, deploying, calling external services that cost money or send messages. The distinction that matters is reversibility, not danger.
Scoped credentials. An agent should have exactly the access the task needs. An agent with production credentials because that was the convenient configuration is a risk with no corresponding benefit.
Isolation for anything substantial. A container or a separate working copy means the worst case is bounded by that environment. This matters more as autonomy increases.
There is a subtler concern worth naming. An agent reads files, issues, comments and documentation, and it acts on what it reads. Content crafted to look like an instruction — in a dependency's documentation, in an issue description, in a code comment — sits in the same context as your task. The boundary between data and instruction is not sharp. Bounding what the agent can do is the durable defence; asking it to be careful is not.
13. Writing a task an agent can execute
Task quality is the input you control that most affects the output.
State the outcome, not just the action. "Add validation to the signup form" is vague. "Reject signups where the email domain is not in the allowed list, returning a field-level error the existing form component can display" is executable.
Name a pattern to follow. "Do this the way the password reset flow does it" is the single highest-value sentence you can include. It gives the agent a concrete example rather than an abstract requirement.
State constraints explicitly. No new dependencies. Do not change the public interface. Keep it in the existing module. Constraints are invisible to an agent unless stated, and violating one is not a mistake it can detect.
Say how it will be verified. Which tests should pass, or what behaviour to check manually. This gives the agent a target and gives you a review criterion.
Include the failing case if there is one. For a bug, the reproduction is worth more than the description. An agent that can reproduce a bug can verify a fix.
Keep the scope to one coherent change. Two related tasks run separately produce better results than one task containing both.
14. Preparing a repository for agents
Some investments pay off disproportionately.
- A fast, reliable test suite. The highest-value item by a wide margin. Everything else is secondary to this.
- A conventions file the agent reads automatically — architecture notes, naming rules, what not to do, how to run things. Most agent tooling looks for one.
- A single documented command to run tests and another to run the build. An agent that has to infer how to test a project wastes actions discovering it.
- Consistent naming. Agents find code by searching for words.
- Examples of each common pattern in the codebase, so "follow the existing approach" has something to point at.
- A clean working tree before starting, so the diff is exactly the agent's work.
Notably, every one of these is good practice independent of AI. Repositories that were already well-maintained get more from agents, which is less a coincidence than an illustration that agents amplify the properties a codebase already has.
15. Reviewing agent output
Review differently from human code, because the failure modes differ.
- Did it solve the stated problem or a nearby one? Check against the task, not just for correctness.
- Did any test file change? Assertions modified to pass are the highest-priority thing to catch.
- Is it consistent with the codebase, or does it introduce a parallel approach?
- Any new dependencies, and were they necessary?
- What happens on failure paths? Usually the thinnest part.
- Is anything over-built relative to the requirement?
- Were files changed that had no business changing? Scope creep in a diff is easy to miss and easy to check.
The practical discipline: review the diff as a whole before reviewing it file by file. Agent diffs fail globally more often than locally, and reading sequentially through nine files makes an inconsistency across them nearly invisible.
16. Where agents genuinely help, and where they do not
Strong fits: bug fixes with a reproduction; mechanical refactoring across many files; adding a feature closely resembling an existing one; writing tests for existing code; dependency upgrades with breaking changes; and codebase-wide renames or migrations.
Weak fits: anything requiring architectural judgement; work depending on context that exists only in people's heads; genuinely novel problems with no similar example; performance work requiring measurement and intuition; and anything where correctness cannot be checked without deep domain knowledge.
The pattern behind the split: agents excel where the target is verifiable and a similar example exists, and struggle where success is a matter of judgement. That is a useful filter to apply before delegating anything.
17. Twelve mistakes
- Vague task descriptions. The agent picks an interpretation and executes it well.
- Delegating without a test suite. Speed on generation, none on verification.
- Not reading the plan. The cheapest possible quality control, skipped.
- Ignoring changed test files. A green suite that verifies nothing.
- Letting it thrash. Repeated failure means re-specify, not wait longer.
- Tasks that are too large. Context exhausts before the careful work happens.
- Starting on a dirty working tree. The diff no longer isolates the agent's work.
- Broad credentials. Risk with no corresponding benefit.
- No conventions file. Every task rediscovers the same rules, or does not.
- Reviewing file by file only. Misses cross-file inconsistency, the characteristic failure.
- Accepting over-built solutions. The agent has no sense of proportion.
- Treating "tests pass" as "task complete". A narrower claim than it sounds.
18. A worked example: a bug fix, step by step
The report: users on a particular subscription tier occasionally see an incorrect renewal date, one day early, and only sometimes.
Task description given to the agent. The symptom, one concrete example with the account identifier and the expected versus actual date, the note that it appears to affect only annual subscriptions, and an instruction to add a regression test. Critically, it also says the fix should not change the public interface of the billing module and should follow the existing date handling in the invoicing code.
Orientation. The agent searches for the term "renewal" and finds eleven files. It reads the subscription model and the renewal calculation. It searches for "annual" and finds a branch handling that case separately. It reads the existing tests for renewal dates, which is where it learns that the codebase distinguishes billing dates from calendar dates — a distinction not obvious from the implementation.
Plan. It proposes: reproduce with a test using the reported example; inspect the date arithmetic in the annual branch; check timezone handling, since "one day early, sometimes" is the signature of a timezone or daylight-saving issue; fix; verify the new test and the existing suite. Reading this plan takes twenty seconds and confirms it understood the problem — which is the point at which a misunderstanding would have been cheap to correct.
Reproduction. It writes a test using the reported account's parameters and runs it. The test passes, which is informative rather than disappointing: the bug is not deterministic from those inputs alone. It then writes a test that iterates over renewal dates across a year and finds failures clustered in late March and late October — the daylight-saving transitions. The hypothesis is now confirmed by evidence rather than assumed.
Investigation. It reads the date arithmetic and finds that the annual branch adds a fixed number of hours rather than incrementing the calendar year, which drifts by an hour across a transition and occasionally crosses midnight. It also reads the invoicing code it was told to follow and finds a helper that does calendar-aware date arithmetic correctly.
Fix. It replaces the arithmetic with a call to that existing helper. Three lines changed in one file. It did not write a new date utility, because the task named an existing pattern and it found it.
Verification. The new test passes across the full year. The existing suite passes except for two tests that were asserting the old, incorrect behaviour — and here the agent does the right thing: rather than modifying them silently, it reports them and asks whether they encode intended behaviour. That question is the correct output, because a test asserting the buggy behaviour could mean either "this test was wrong" or "you have misunderstood the requirement", and the agent cannot distinguish those.
Review. The diff is three lines of fix, one new test file and a question about two existing tests. Reviewing it takes five minutes. The two tests turn out to have been written to match the buggy output during an earlier investigation, so they are corrected.
What made this work. A reproduction case in the task description gave the agent a target it could verify against. Naming an existing pattern to follow prevented it inventing a parallel date utility. A fast test suite let it iterate — it ran the tests eleven times over the task. And the constraint about not changing the public interface kept a three-line fix from becoming a refactor.
How the same task fails. Given only "renewal dates are sometimes wrong, please fix", the same agent searches, finds the annual branch, notices the arithmetic looks fragile, rewrites it with its own date utility, adds configuration for timezone handling nobody asked for, and produces a two-hundred-line diff that passes the tests and introduces a second way of handling dates. Every step is defensible. The result is worse than no change, and the only difference between the two runs is four sentences in the task description.
19. Frequently asked questions
How much autonomy should I give an agent?
As much as your verification allows. In a repository with a fast, trustworthy test suite and clean version control, a long autonomous run is reasonable because the worst case is a diff you discard. Without those, keep tasks small and review each one, because you have no faster way to establish correctness than reading everything.
Why does it sometimes go in circles?
Almost always because it has misunderstood something fundamental — the wrong file, the wrong layer, or a misread requirement — and further iteration cannot fix a wrong premise. Thrashing is a signal to stop and re-specify rather than to let it continue. Reading its plan before it starts prevents most instances.
Will it modify my tests to make them pass?
Sometimes, and it is not deception — the objective was stated as making the tests pass. It is the most important thing to check in any agent diff. Treat any change to an assertion as requiring an explicit justification, and prefer agents that report conflicting tests rather than silently adjusting them.
What is the highest-value preparation I can do?
A fast, reliable test suite, by a wide margin. It is the mechanism by which an agent verifies its own work, and every other improvement is secondary to it. After that, a conventions file the agent reads automatically and consistent naming across the codebase.
Can an agent do something destructive?
It can run whatever commands the harness permits, so the defence is the environment rather than the instructions. Start from a clean, committed working tree; require approval for irreversible actions; scope credentials to the task; and isolate anything substantial. With those, the worst realistic outcome is changes you discard.
How large a task can it handle?
Technically quite large; practically, smaller is better. Everything the agent reads occupies context, and a broad task exhausts that room before the careful work happens. Two focused tasks run separately routinely beat one task containing both, and each is independently reviewable.
Should I read the plan before it starts?
For anything non-trivial, yes. It takes under a minute and it catches the most common failure — the agent executing a reasonable interpretation of an ambiguous description. Catching that before a four-hundred-line diff exists is the cheapest quality control in the whole workflow.
Do agents make senior developers or junior ones more productive?
Senior, generally, because the skills that matter are specifying tasks precisely and reviewing output critically — both of which come from experience. Juniors benefit too, with a caveat: the tasks they traditionally learned on are exactly the ones most easily delegated, so it is worth requiring them to explain what they accepted.
Key takeaways
- The loop is the innovation — acting, observing, and deciding again with the result in hand.
- The harness executes, not the model. Every permission decision belongs there.
- Tests determine quality more than the model does, because they close the loop.
- Name an existing pattern to follow. The single highest-value sentence in a task description.
- Read the plan before the diff. Catches misunderstandings while they are still cheap.
- Check whether test files changed. A passing suite that was edited to pass verifies nothing.
Coding agents are not magic and not autocomplete. They are a model placed in a loop with the ability to check its own work, and everything about using them well follows from that — give them something to check against, something to imitate, and a task specific enough that there is only one reasonable reading.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch