Skip to main content
Blog

Test-Driven Development: Writing Better Code Through Tests

Last updated Development

Test-driven development is the most misunderstood practice in software engineering. Its critics describe it as writing tests first, which sounds like bureaucracy with extra steps. Its advocates describe it as a design technique, which sounds like mysticism. Both are pointing at the same thing from different angles: TDD is a feedback loop so tight that it changes the shape of the code you write, and the tests are a by-product rather than the goal.

This guide explains the practice properly — the loop, what each step is actually for, how to choose what to test, how to avoid the traps that make people abandon it, and where it genuinely does not apply. It assumes you can write code but no prior testing discipline, and it uses plain examples throughout.


What you will learn
  • What the red-green-refactor loop does beyond producing tests
  • How to choose the right first test, and the right size of step
  • The difference between testing behaviour and testing implementation
  • Test doubles explained without the terminology fog
  • How TDD changes the structure of code, and why that is the real benefit
  • Where TDD works badly, and what to do instead
In this article
  1. What TDD actually is
  2. The loop, step by step
  3. A worked example
  4. Choosing the first test
  5. Behaviour, not implementation
  6. Test doubles without the jargon
  7. The shape of the suite
  8. How TDD changes design
  9. Refactoring: the step everyone skips
  10. Legacy code: TDD when there are no tests
  11. Where TDD works badly
  12. Ten anti-patterns
  13. Measuring whether it is working
  14. Adopting TDD in a team that has never done it
  15. Frequently asked questions

1. What TDD actually is

TDD is a cycle of three steps repeated in periods of a few minutes: write a failing test for the next small piece of behaviour, write the simplest code that makes it pass, then improve the design while the tests keep you honest. That is the whole rule.

What makes it powerful is not the tests. It is the constraint. Writing the test first forces you to answer three questions before you have the comfort of an implementation to hide behind:

  • What does this thing do? You must state the behaviour in concrete terms, with real inputs and real expected outputs. Vague requirements become obviously vague.
  • How will it be called? You design the interface from the caller's perspective, before you know how it will work inside. This is the single largest design benefit, and it is the reason TDD-written code tends to be easier to use.
  • How will I know it is done? The test defines the finish line. Without it, "done" drifts, and you keep polishing code that already worked.
A useful reframing: TDD is not a testing practice that happens to improve design. It is a design practice that happens to leave behind a regression suite.

2. The loop, step by step

Red — write a failing test

Write a test for behaviour that does not exist yet, and run it. It must fail, and it must fail for the reason you expect. A test that passes immediately is telling you something: either the behaviour already exists, or the test does not actually assert anything meaningful. Both are worth knowing before you continue.

Keep the test small. One behaviour, one reason to fail. If you find yourself writing three assertions about unrelated things, you are describing three tests.

Green — make it pass, simply

Write the least code that turns the test green. Not the most elegant, not the most general — the least. This feels wrong to experienced developers, and it is the step where most people cheat by writing the full implementation they had already imagined.

The discipline has a purpose. Writing minimal code keeps the step small enough that when something breaks, you know exactly which change caused it. It also exposes over-engineering: if the simple version passes every test you can think of, the elaborate version you were about to write was speculation.

Refactor — improve the design

Now, with a passing test as a safety net, improve the code. Remove duplication, rename things to what they actually mean, extract a function that has emerged, simplify a conditional. Run the tests after every change.

This is the step that gets skipped, and skipping it turns TDD into "write tests first and accumulate mess" — which is worse than not doing TDD at all, because you now have tests pinning bad design in place.

The rhythm

Each full cycle should take somewhere between thirty seconds and a few minutes. If a cycle takes twenty minutes, the step was too big; break the behaviour into smaller pieces. The tight rhythm is not stylistic preference — it is what keeps the debugging surface small, because the only thing that changed since the last known-good state is a handful of lines.

3. A worked example

Take a familiar problem: a shopping basket that applies a discount when the subtotal exceeds a threshold. What matters is the sequence of cycles rather than the finished result — each row below is one complete loop, taking a minute or two.

CycleRed — the failing testGreen — the least codeRefactor
1An empty basket totals zeroThe total method returns zero, unconditionally. Deliberately trivial, and correct for everything asserted so far.Nothing to do
2A basket with one item totals that item's priceStore items in a list; total sums their prices. The unconditional zero is now clearly wrong, so the test forced the real implementation.Nothing yet — the code is already minimal
3Ten percent comes off when the subtotal exceeds the thresholdA single conditional applied after the sumThe discount rule is now a distinct concept and a candidate for its own object the moment a second rule appears
4What happens at exactly the threshold?Whichever behaviour the team decides — but the decision has now been made explicitlyRename the threshold constant to state the rule it encodes

Notice what has happened without anyone designing it. The basket exposes exactly two operations, because those are the only two any test needed. Prices ended up as whole minor units rather than decimals, because cycle three forced a rounding decision early rather than at the point where a customer notices a penny discrepancy. And the discount arrived as its own cycle, which is why it is visible as a separate concept rather than buried inside a summation.

Cycle four is the interesting one. Writing tests first tends to surface boundary questions while they are still cheap to answer, because you cannot write the test without first deciding what the answer should be. Written afterwards, that boundary is usually whatever the implementation happened to do.

4. Choosing the first test

The hardest part of TDD for beginners is knowing where to start. Three heuristics cover most situations.

Start with the degenerate case

Empty list, zero items, null input, no matches. These tests are trivial to write, they establish the interface, and they often reveal a decision you had not made — what should happen when there is nothing?

Then take the simplest meaningful case

One item, one user, one rule. This is where the real shape appears, and it usually takes the implementation most of the way.

Then attack the boundaries

The values on either side of a threshold, the maximum length, the moment a limit is exceeded. Boundaries are where defects cluster in every codebase ever measured, and they are exactly what people forget to test when writing tests afterwards.

What to avoid: starting with the most complex case. It produces a large step, a long red phase, and a temptation to write the entire implementation at once — which is TDD in name only.

5. Behaviour, not implementation

This single distinction determines whether your test suite becomes an asset or a liability.

A test coupled to behaviour says: given this input, the observable result is that. It survives refactoring, because refactoring by definition does not change behaviour. A test coupled to implementation says: this method calls that method with these arguments. It breaks the moment you restructure anything, even when nothing observable changed.

Implementation-coupledBehaviour-coupled
Asserts that a private helper was calledAsserts the returned value is correct
Verifies the order of internal callsVerifies the outcome the caller sees
Reaches into internal state to check a fieldChecks via the public interface
Mocks every collaboratorMocks only what crosses a real boundary
Breaks on every refactorBreaks only when behaviour changes

The practical test: if you rewrite the implementation completely while keeping the same external contract, should the test still pass? If yes, it is a behaviour test. If no, you have written a test that documents your current code rather than your intended behaviour — and its only effect will be to make future improvements more expensive.

This is also why a suite with very high coverage can still be nearly worthless. Coverage measures which lines ran, not whether anything meaningful was asserted about them.

6. Test doubles without the jargon

Sometimes a unit depends on something you cannot or should not use in a test: a payment gateway, the clock, a database, an email service. You substitute a stand-in. The terminology is unnecessarily confusing, so here it is plainly.

KindWhat it doesUse when
StubReturns canned answersThe dependency supplies data you need
FakeA real, simplified implementation — an in-memory repositoryYou need realistic behaviour without the real system
SpyRecords what happened for later inspectionYou need to confirm an effect occurred
MockAsserts that specific calls were madeThe interaction itself is the behaviour

Two rules prevent nearly all test-double pain. Double at boundaries you own — your own interface for sending email, not the third-party library's class. When the vendor changes, your tests do not. And prefer fakes to mocks: an in-memory repository produces tests that read like descriptions of behaviour, while a wall of mock expectations produces tests that read like a transcript of the implementation.

The strongest signal that you are over-doubling is a test that requires five mocks to set up. That is not a testing problem; it is the design telling you the unit has too many collaborators.

7. The shape of the suite

Not every test should be written the same way. The useful mental model is a pyramid, and its proportions matter.

LevelScopeSpeedShare of suite
UnitOne class or function, doubles at boundariesMillisecondsMost of it
IntegrationReal database, real HTTP, one processTens to hundreds of msA meaningful minority
ContractTwo services agree on a message shapeFastOne per boundary
End to endThe whole system through the real interfaceSecondsA handful only

Inverting this — few unit tests, many end-to-end tests — produces the ice-cream cone: a slow, flaky suite that people re-run until it goes green. Once a team learns to re-run failures, the suite has stopped providing information, and every subsequent investment in it is wasted.

The end-to-end tests you keep should cover only the journeys where failure is unacceptable: sign in, pay, the one action the product exists to perform. Everything else is better tested lower down, where failures point directly at a cause.

8. How TDD changes design

The claim that TDD improves design sounds like faith until you notice the mechanism, which is entirely mundane: code that is hard to test is hard to test for specific, diagnosable reasons, and fixing those reasons produces better structure.

Testing difficultyWhat it revealsThe design fix
Needs a lot of setupToo many dependenciesSplit responsibilities, or pass what is needed
Cannot control the resultHidden dependency on time, randomness or environmentInject the clock or generator
Must check a database to verifyBusiness logic mixed with persistenceSeparate decision from effect
Requires many mocksThe unit orchestrates too muchIntroduce a smaller collaborator with a real interface
Test name needs "and"The unit does two thingsSplit it

The recurring pattern in all of these is separating decisions from effects. A function that calculates what should happen is trivially testable — inputs in, result out. A function that performs an effect is trivially testable if it does nothing else. A function that decides and acts in the same breath requires elaborate machinery to test, which is the design smell TDD surfaces continuously.

This is why teams that adopt TDD tend to drift towards a similar architecture without being told to: a core of pure decision-making logic, surrounded by a thin shell that talks to the outside world. Nobody mandated it; the tests made every other arrangement uncomfortable.

9. Refactoring: the step everyone skips

Refactoring means changing the structure of code without changing its behaviour. The tests are what make it safe, and the tests are also what make it feel unnecessary — everything is green, so why touch it?

Because the green code from the previous step is deliberately the simplest thing that worked, which means it is frequently duplicated, poorly named, or holding a shape that no longer matches what the code does. Small, continuous cleanup is what prevents the accumulation that eventually gets called "we need a rewrite."

Useful moves, roughly in order of how often they apply:

  • Rename to what the thing actually means now, not what it meant when you created it.
  • Extract a function when a block needs a comment to explain itself — the comment is usually the function name.
  • Inline an indirection that no longer earns its keep.
  • Remove duplication, but only on the third occurrence; two similar things are often coincidence, and premature unification couples things that later need to diverge.
  • Replace a conditional with polymorphism when the same branch structure appears in several places.

The rule that makes it safe: refactor only when the tests are green, change one thing at a time, and run the tests after each change. If you find yourself refactoring with a failing test, stop — you are now doing two things at once and cannot tell which one broke.

10. Legacy code: TDD when there are no tests

Most real work happens in code that was written without tests, and the standard advice to "write tests first" is unhelpful when the code cannot be instantiated without a database, a queue and three environment variables. The practical sequence is different.

  1. Write a characterisation test. Not a test of what the code should do — a test of what it does. Call it, observe the output, assert that output. This is a safety net, not a specification.
  2. Find a seam. A place where behaviour can be substituted without editing the code you are afraid of: a constructor parameter, a method you can override, a configuration value. Seams are what make untestable code testable.
  3. Break the dependency minimally. Extract the offending call into a method you can override, or introduce a parameter with a default. This is the one edit you make without test coverage, so keep it mechanical and small.
  4. Now TDD the change. With a seam and a characterisation test, you can write a failing test for the new behaviour and proceed normally.

The governing principle is to leave code better than you found it in the area you touched, without attempting to fix everything. Wholesale test-retrofitting projects are almost always cancelled halfway; incremental improvement around actual changes compounds quietly and never needs approval.

11. Where TDD works badly

Honest advocacy requires naming the limits.

  • Exploratory work. When you do not know what you are building — trying an API, prototyping an interaction — tests slow discovery. Explore freely, then delete the spike and rebuild it test-first with what you learned. The mistake is keeping the spike.
  • User interface layout. Whether something looks right is not expressible as an assertion. Test the logic behind the screen, and use visual review or snapshot comparison for appearance.
  • Genuinely emergent behaviour. Machine-learning outputs, physics simulations and heuristics have no single correct answer. Test properties and invariants instead: outputs are in range, results are stable across runs, known-bad inputs are rejected.
  • Thin glue code. A function that maps one data shape onto another with no logic gains little from a test. Cover it once through an integration test and move on.
  • Extreme deadline pressure with a throwaway artefact. A genuine one-week prototype that will be deleted does not need a regression suite. The trap is that such prototypes are frequently not deleted.

12. Ten anti-patterns

  1. Writing tests after the code and calling it TDD. Perfectly valid practice, different benefits — you get regression safety but none of the design pressure.
  2. Skipping refactor. Produces tested mess, which is harder to improve than untested mess.
  3. Testing private methods. If it needs its own test, it wants to be a separate unit with a public interface.
  4. One test with fifteen assertions. When it fails you learn one thing instead of fifteen.
  5. Shared mutable fixtures. Tests that pass alone and fail together, or pass in one order and fail in another. Build fresh state per test.
  6. Asserting on log output. Logs are for humans; tests that depend on their wording break constantly.
  7. Chasing a coverage number. Produces tests written to touch lines rather than to check behaviour.
  8. Sleeping in tests. A fixed delay is both slow and unreliable. Control time explicitly or wait on a condition.
  9. Tolerating flaky tests. One tolerated flake teaches the team that red does not mean broken, and the suite's value collapses.
  10. Mirroring the implementation in test structure. One test class per production class, mocking every collaborator, means every refactor is a test rewrite.

13. Measuring whether it is working

The point of the practice is not test count. Watch these instead:

SignalHealthyWarning
Suite runtimeUnit tests under a minutePeople stop running them locally
Flake rateEffectively zeroAny tolerated flake
Defect escape rateTrending downBugs found in production that a unit test could have caught
Change confidenceRefactoring happens routinelyPeople avoid touching certain files
Test churnTests change when behaviour changesTests rewritten on every refactor
Time to diagnoseA failing test names the causeA failure requires debugging to interpret

The last row is the most telling. In a healthy suite, the name of the failing test tells you what broke before you read a line of the diff.

14. Adopting TDD in a team that has never done it

Individual adoption is a matter of practice. Team adoption is a matter of removing the conditions that make the practice impossible, and those conditions are usually environmental rather than attitudinal.

Make the suite fast before asking anyone to run it constantly. The loop depends on getting feedback in seconds. If the unit suite takes four minutes, nobody will run it between edits, and TDD degenerates into writing tests before a long build. Profile the slowest tests first — in most codebases a handful of tests that quietly touch the network or the filesystem account for most of the runtime.

Make it possible to run one test. Sounds trivial; frequently is not. If running a single test requires starting three services and seeding a database, the practice is dead before it starts. Investing a day in a test harness that spins up isolated state per test pays for itself within a fortnight.

Start where the value is undeniable. Bug fixes, as noted earlier, are the ideal entry point, because writing a failing test that reproduces a defect is something nobody argues with. The second easiest entry point is a new module with no legacy entanglement — a fresh piece of pure logic where the practice can be demonstrated without fighting the existing architecture.

Pair on the first few sessions. The most common reason people abandon TDD in week one is that they get stuck deciding what the first test should be, conclude the practice is impractical, and go back to what they know. Twenty minutes of pairing at the right moment prevents that conclusion, and it transfers the rhythm — the sense of how big a step should be — far more effectively than any written explanation.

Do not mandate it. Mandated TDD produces compliance artefacts: tests written after the code and reordered in the commit, or tests that assert nothing meaningful in order to satisfy a reviewer. What works is making the practice easy, demonstrating it on real work, and letting the reduction in debugging time do the persuading. Teams that adopt it voluntarily keep doing it; teams that are ordered to stop the moment the pressure lifts.

Expect a dip. The first month is slower, and the team should be told so explicitly, because an unexpected dip gets attributed to the practice being wrong rather than to the practice being new. The return arrives when the first significant refactor happens without anyone being afraid, which is usually somewhere between six weeks and three months in.

15. Frequently asked questions

Does TDD slow development down?

It slows the first pass and speeds everything afterwards. Studies and experience broadly agree that initial implementation takes somewhat longer while defect rates fall substantially. Since most of a system's cost is incurred after the first version — changing it, debugging it, being afraid of it — the trade is usually favourable. It is genuinely unfavourable for code that will be written once and never modified, which is rarer than people assume.

What coverage percentage should we target?

None, as a target. Coverage is a diagnostic, not a goal — it tells you which code has no tests at all, which is useful, but a mandated number produces tests written to satisfy the number. Look instead at whether the code you are afraid to change is covered by meaningful behaviour tests. That question has no percentage answer and is far more informative.

How is TDD different from BDD?

Mostly vocabulary and audience. Behaviour-driven development uses scenario language that non-developers can read, and typically operates at a higher level — describing a user journey rather than a function. The loop is identical. Many teams use both: scenarios at the feature boundary, unit-level TDD inside. They are complementary rather than competing.

Should tests use mocks or real dependencies?

Real ones wherever they are fast and deterministic. An in-memory database, a local queue or a fake HTTP server gives you realistic behaviour with none of the brittleness of mocking every call. Reserve mocks for genuinely external systems and for cases where the interaction itself is the behaviour under test, such as verifying that a notification was sent exactly once.

What do I do when the test is hard to write?

Treat it as information rather than an obstacle. Difficulty in testing almost always identifies a specific structural problem — hidden dependencies, mixed responsibilities, or logic entangled with input and output. Ask what would make this easy to test, and you usually find that the answer is also a better design. This is precisely the feedback the practice exists to give you.

Can TDD work with a database-heavy application?

Yes, with a separation. Business rules are tested as pure logic with no database. Repository and query behaviour is tested against a real database instance — modern container tooling makes this fast enough to run routinely. What does not work is trying to unit-test SQL by mocking a database driver: you end up asserting that your code produces the query string you wrote, which proves nothing about whether it returns the right rows.

How small should a step be?

Small enough that if the test fails unexpectedly, you can find the cause without a debugger. In practice that means a few lines of production code per cycle when the territory is unfamiliar, and larger steps when it is routine. Experienced practitioners vary step size continuously — shrinking it when surprised, growing it when confident. Struggling for more than a few minutes is the signal to take a smaller step.

Our team is convinced this is not worth it. What is the smallest useful start?

Apply it only to bug fixes. Write a failing test that reproduces the bug, fix it, keep the test. This is uncontroversial — nobody argues against reproducing a defect before fixing it — it produces immediate value, and it builds the habit on the code most likely to break again. Teams that start here often extend the practice on their own, because the loop is more pleasant than the alternative once it is familiar.

Key takeaways

  • TDD is a design practice. The regression suite is a valuable by-product, not the purpose.
  • Keep cycles measured in minutes. Small steps keep the debugging surface small.
  • Test behaviour, never implementation. A test that breaks on every refactor is a liability.
  • Do not skip refactor. Tested mess is harder to fix than untested mess.
  • Difficulty testing is a design signal. Ask what would make it easy, then do that.
  • Prefer fakes to mock walls. Five mocks in one test is the design speaking to you.

The practice is simple enough to describe in three words and takes months to internalise, because the hard part is not writing tests — it is resisting the urge to write more code than the current test demands. That restraint is what produces code that is small, well-named, and safe to change.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch