Microservices solve an organisational problem with a technical mechanism. That sentence explains both why they work and why so many teams regret adopting them. If your constraint is that too many people are trying to change one codebase and release it together, splitting into independently deployable services genuinely helps. If your constraint is anything else, you have bought distributed systems complexity and paid for it with a problem you did not have.
This guide covers what microservices actually require: how to find boundaries that hold, how services should communicate, why data ownership is the hardest part, what fails in distributed systems and how to survive it, and what operational maturity the architecture assumes. It also covers, honestly, when not to do this.
What you will learn
- What problem microservices solve, and the prerequisites they assume
- How to find service boundaries that do not need constant renegotiation
- Synchronous versus asynchronous communication, and when each is right
- Why shared databases break everything, and the patterns that replace transactions
- Resilience patterns: timeouts, retries, circuit breakers, idempotency
- Observability, deployment and the operational cost of distribution
- The problem microservices solve
- The prerequisites nobody mentions
- Finding boundaries
- Service granularity
- Communication patterns
- Data ownership
- Consistency without distributed transactions
- Resilience in a distributed system
- The API gateway and edge concerns
- Observability
- Deployment and versioning
- Testing across services
- Migrating from a monolith
- Twelve failure patterns
- A worked example: extracting the first service
- Frequently asked questions
1. The problem microservices solve
A monolith becomes painful for reasons that are mostly about people rather than technology. Fifteen teams merging into one codebase produce constant conflicts. One team's bug blocks everyone's release. A change to a shared module requires coordinating with six other teams. The release train runs monthly because that is the only cadence at which everyone can be synchronised.
Microservices address this by making deployment independent. Each service has its own codebase, pipeline, database and release schedule. Teams stop coordinating and start integrating through contracts.
Two secondary benefits are real but smaller than commonly claimed. Independent scaling matters when one component has a dramatically different load profile — a search service handling ten times the traffic of the rest. Technology diversity lets a team choose a different language where it genuinely helps, though in practice most organisations discover that standardisation is worth more than choice.
What microservices do not do: make the system faster (network calls are slower than function calls), simpler (they add failure modes), or cheaper (operational overhead multiplies). Anyone promising those is describing a different architecture.
The clarifying question: can two teams currently deploy independently? If yes, you do not have the problem microservices solve. If no, ask whether the obstacle is genuinely the codebase, or whether it is a shared database, a manual release process, or an approval queue — because splitting services fixes none of those.
2. The prerequisites nobody mentions
Microservices assume operational capabilities that a monolith does not require. Adopting the architecture without them produces a distributed monolith: all the complexity, none of the independence.
| Capability | Why it becomes mandatory |
|---|---|
| Automated deployment | Twenty services deployed by hand is twenty times the manual work |
| Distributed tracing | A request crossing eight services cannot be debugged from logs alone |
| Centralised logging | Nobody can log into twenty machines to find one error |
| Service discovery | Hard-coded addresses fail the first time an instance moves |
| Contract testing | Otherwise integration bugs are found in production |
| On-call ownership | Each service needs someone who answers when it breaks |
| Infrastructure as code | Provisioning per service manually does not scale past a handful |
The honest sequencing advice: build these capabilities first, inside a monolith, where the cost of getting them wrong is low. A team with automated deployment, good observability and clear module boundaries can split services when it needs to. A team without them will find that splitting makes every existing weakness considerably worse.
3. Finding boundaries
The single decision that determines whether a microservices architecture is pleasant or miserable is where the boundaries fall. Bad boundaries mean every feature requires changing four services, which reintroduces coordination while keeping the network calls.
Boundaries follow the business, not the technology
The failure pattern is splitting by technical layer — an API service, a business logic service, a data service. That guarantees every change crosses all three. Split instead by business capability: ordering, inventory, payments, shipping, identity. Each owns a coherent area that changes for its own reasons.
The language test
Listen to how the business talks. Where the same word means different things to different teams, you have found a boundary. "Customer" in sales means a prospect with a pipeline stage. "Customer" in billing means an entity with a payment method. "Customer" in support means someone with a ticket history. Those are three different concepts that share a name, and forcing them into one model produces an object with forty fields that nobody fully understands.
The change test
Look at your commit history. Which files change together? Things that consistently change together belong together, regardless of how the architecture diagram groups them. This is the most empirical guidance available and is routinely ignored in favour of theoretical decomposition.
The team test
A service should be owned by one team. Two teams sharing a service reintroduces exactly the coordination microservices were meant to remove. Conversely, one team owning fifteen services is usually a sign of over-decomposition — nobody can hold that many operational concerns at once.
4. Service granularity
"Micro" is misleading. There is no correct size, but there are recognisable symptoms of getting it wrong in either direction.
| Too large | Too small |
|---|---|
| Multiple teams contend for the same repository | Most changes require deploying several services together |
| Deployments are risky because the blast radius is wide | Simple operations need three network calls |
| Different parts have very different scaling needs | More time spent on plumbing than on features |
| The codebase is hard to hold in one's head | Every service has its own pipeline nobody maintains |
| Unrelated concerns share a release schedule | Debugging requires tracing through many hops |
A workable heuristic: a service should be small enough that one team can understand it fully and rewrite it in a few weeks if necessary, and large enough that most changes to its business capability stay inside it. In practice that produces fewer, larger services than the word "micro" suggests — and teams that err towards larger services almost always report fewer regrets than those that err smaller.
5. Communication patterns
Synchronous request and response
One service calls another and waits. Simple, familiar, easy to reason about, and it introduces temporal coupling: the caller cannot succeed unless the callee is available right now. Chains of synchronous calls multiply this — five services at ninety-nine percent availability each give roughly ninety-five percent combined, and every one adds its latency to the total.
Appropriate when the caller genuinely needs the answer to proceed: fetching data to render a response, validating something before accepting a request.
Asynchronous messaging
One service publishes an event; others consume it when they can. The publisher does not know or care who is listening, which removes both temporal and logical coupling. The consumer can be down for an hour and catch up afterwards.
Appropriate for anything that does not need an immediate answer: notifying other systems that something happened, triggering downstream processing, updating read models. The cost is eventual consistency — for a period, different services have different views of the world — and debugging that is genuinely harder because there is no single call stack.
Choosing between them
Default to asynchronous for anything that is a notification of fact ("an order was placed") and synchronous for anything that is a question ("is this card valid"). The most common architectural mistake is using synchronous calls for notifications, which couples the caller's availability to systems it should not depend on at all.
Event design
Two styles, and the difference matters. A thin event carries only an identifier, requiring consumers to call back for details — which reintroduces coupling and load. A fat event carries the data consumers need, letting them work independently but making the schema a contract you must version carefully. Most systems settle on fat events containing the fields consumers actually need, with a link for anything unusual.
6. Data ownership
This is where microservices projects most often go wrong, and it is the rule that cannot be bent: each service owns its data exclusively. No other service reads its tables directly.
Teams resist this because a shared database is convenient. Joining across domains in one query is easy; making a network call and joining in memory is not. But a shared database defeats the entire purpose. If two services read the same table, neither can change its schema without coordinating — which is precisely the coupling you were trying to remove, now with added network latency.
Living without joins
Three patterns replace the cross-domain join:
- Composition at the edge. The caller fetches from each service and combines. Simple, and appropriate when the number of calls is small and bounded.
- Data duplication via events. A service keeps a local copy of the fields it needs from another domain, updated by subscribing to events. This is the workhorse pattern, and the duplication is deliberate rather than a smell — the copy is a cache with an update mechanism.
- A dedicated read model. A separate store built specifically for a query pattern, populated from events across several services. Appropriate for reporting, search and dashboards that genuinely span domains.
The instinct that duplication is wrong comes from normalised database design, where a single system owns everything. In a distributed system, the alternative to duplication is coupling, and coupling is worse.
7. Consistency without distributed transactions
In a monolith, a database transaction makes several changes atomic. Across services, that guarantee is unavailable in any practical form, and attempting to recreate it via two-phase commit produces a system that is both slow and fragile.
The pattern that works is the saga: a sequence of local transactions, each publishing an event that triggers the next, with a defined compensating action for each step in case a later one fails. An order that reserves stock, takes payment and schedules shipment does each locally; if payment fails, a compensating action releases the stock.
Two important consequences. First, compensation is not rollback — the intermediate state was visible, and undoing it may require a business decision rather than a database operation. Refunding a payment is not the same as it never having happened. Second, the system is eventually consistent by design, so interfaces must be built to show pending states honestly rather than pretending everything is immediate.
One implementation detail solves a surprising number of bugs: the transactional outbox. Writing to your database and publishing an event are two operations that can fail independently, leaving state changed but nobody notified. Instead, write the event to a table in the same transaction as the state change, and have a separate process publish from that table. The two operations become one atomic write, and delivery becomes an at-least-once guarantee that consumers handle with idempotency.
8. Resilience in a distributed system
Distribution introduces failure modes that do not exist in a single process. A function call either returns or throws; a network call can also hang forever, succeed after the caller gave up, or succeed on the server while the response is lost.
| Pattern | What it prevents | Getting it wrong |
|---|---|---|
| Timeouts | Threads waiting forever on a dead dependency | No timeout at all, or one longer than the caller's own |
| Retries with backoff and jitter | Transient failures becoming user-visible | Immediate synchronised retries that amplify an overload |
| Circuit breakers | Hammering a service that is already struggling | Thresholds so loose the breaker never opens |
| Bulkheads | One slow dependency exhausting all resources | A shared connection pool for every downstream call |
| Idempotency | Retries causing duplicate side effects | Assuming a timeout means the request did not arrive |
| Graceful degradation | Total failure when one non-essential part is down | Treating every dependency as mandatory |
Two rules deserve emphasis. Only retry idempotent operations, or make operations idempotent with a client-supplied key — otherwise retrying a payment charges twice. And set timeouts that decrease down the call chain: if the caller times out at three seconds, a downstream call with a five-second timeout is pointless work that keeps resources occupied after nobody is listening.
The most damaging distributed failure is the cascade: one slow service causes callers to accumulate waiting requests, which exhausts their capacity, which slows their callers, until the whole system is unavailable because one component became slightly slow. Timeouts, bulkheads and circuit breakers exist specifically to break that chain, and they must be in place before the incident rather than added after it.
9. The API gateway and edge concerns
Clients should not call twenty services directly. A gateway at the edge handles the concerns that would otherwise be duplicated everywhere: authentication, rate limiting, request routing, and translation between the shape clients want and the services that exist.
Keep the gateway thin. The recurring failure is a gateway that accumulates business logic until it becomes a monolith with a routing table — at which point every team must change it, and the coordination problem returns at the edge.
Where different client types need genuinely different shapes — a mobile app wanting fewer, richer responses than a web application — a backend-for-frontend per client type is often cleaner than one gateway trying to serve both. Each is owned by the team that owns that client, which keeps the ownership boundary intact.
10. Observability
In a monolith, a stack trace tells you what happened. In a distributed system, the story is spread across services, and reconstructing it requires deliberate instrumentation.
Distributed tracing is not optional. Every request gets an identifier at the edge, propagated through every call and included in every log line. Without it, debugging a slow request means correlating timestamps across services by hand, which does not scale past about three services.
Structured logs shipped centrally, with the trace identifier, service name and version on every entry. Free-text logs scattered across machines are effectively write-only.
Metrics per service: request rate, error rate, latency distribution and saturation. These four answer most operational questions and are worth standardising so every service reports them identically.
Service dependency mapping, ideally generated from traces rather than maintained by hand. When something breaks, the first question is what else is affected, and a hand-drawn diagram is always out of date.
Alert on user-facing symptoms rather than on individual service health. A single instance being unhealthy is normal; the error rate on the checkout journey rising is not. Alerting per service produces noise that trains people to ignore alerts.
11. Deployment and versioning
Independent deployment is the entire point, which means services must tolerate their dependencies being at different versions — permanently, not just briefly during a rollout.
The governing rule is backward-compatible change only. Add fields, never remove or repurpose them. Accept both old and new message formats during a transition. Never require a coordinated deployment of two services, because the moment you do, you have a distributed monolith and every claimed benefit evaporates.
The safe sequence for a breaking change is expand, migrate, contract: add the new form alongside the old, move consumers over one at a time, verify nobody uses the old form, then remove it. It takes longer and it is the only approach that preserves independence.
Each service needs its own pipeline, its own version, and its own rollback path. Shared release trains reintroduce coordination; if all services deploy together on Thursday, you have a monolith distributed across a network.
12. Testing across services
Testing distributed systems requires rebalancing where confidence comes from, because full end-to-end tests are slow, flaky and require every service running.
- Unit tests within each service, as usual, forming the bulk of the suite.
- Contract tests at each boundary. The consumer declares what it expects; the provider verifies it can satisfy that. Both run independently, which catches integration breaks without either team running the other's code.
- Component tests exercising a single service with its dependencies stubbed, verifying it behaves correctly in isolation including its failure handling.
- A small number of end-to-end journeys covering only what must never break.
- Production verification: synthetic transactions running continuously against production, which catch problems no test environment reproduces.
Contract testing is the highest-value practice in this list and the most often skipped. It replaces the integration environment where everything is deployed together — an environment that is always broken, always out of date, and always the bottleneck.
13. Migrating from a monolith
Big-bang rewrites fail with remarkable consistency. The pattern that works is incremental extraction, with the monolith running throughout.
- Modularise in place first. Establish clear internal boundaries inside the monolith, with modules communicating through defined interfaces rather than reaching into each other. This is where you discover whether your proposed boundaries are right, and it is far cheaper to fix them here.
- Extract the edges first. Choose a module with few dependencies and clear ownership — notifications, document generation, search. Prove the operational pattern on something whose failure is survivable.
- Route through a facade. Put an interception layer in front, so traffic can be directed to the old or new implementation without callers knowing. This is what makes each step reversible.
- Run both and compare. Send traffic to both implementations, use the old result, and log differences. This finds behaviour nobody documented, which every long-lived monolith contains.
- Move the data last. Keep the new service reading the old database initially, then migrate ownership once behaviour is proven. Splitting code and data simultaneously doubles the risk of every step.
- Delete the old path. The step most often skipped, leaving two implementations forever. Removal is what realises the benefit.
Expect the first extraction to take considerably longer than estimated, because it includes building the operational capability the architecture needs. The second is much faster. If it is not, the boundaries are wrong.
14. Twelve failure patterns
- Splitting by technical layer. Every feature crosses every service.
- A shared database. Nobody can change a schema; the coupling never went away.
- Synchronous chains. Availability multiplies down and latency adds up.
- Coordinated releases. A distributed monolith with extra network calls.
- Breaking changes without a transition. One deployment takes down its consumers.
- Retries without idempotency. Duplicate charges and duplicate records.
- No timeouts. One slow dependency exhausts every caller.
- No distributed tracing. Debugging becomes archaeology across log files.
- A gateway that grew business logic. The coordination bottleneck reappears at the edge.
- Services without owners. Nothing is patched, nothing is on call, nothing is deleted.
- Over-decomposition. More plumbing than product.
- Adopting it without the operational prerequisites. Every existing weakness multiplied by the number of services.
15. A worked example: extracting the first service
Consider a retail platform that has grown into a single application maintained by five teams. Releases happen fortnightly because that is the only cadence at which everyone can be coordinated, and a bug in the reporting module has twice delayed a checkout fix. That is a genuine microservices problem: the teams need independent deployment, and the codebase is what prevents it.
The first decision is what to extract, and the answer is not the most important thing. Checkout is the highest-value area and the worst possible first candidate, because a mistake there is a revenue incident. The team instead chooses notifications — email and message sending — which has few inbound dependencies, a clear owner, and a failure mode that is annoying rather than catastrophic. The purpose of the first extraction is to build operational capability, and that is best done where mistakes are survivable.
Before any code moves, the module is isolated inside the monolith. Every call to the notification code is routed through a single interface rather than scattered across twenty files. This takes two weeks and produces something valuable on its own: it reveals that four parts of the system were reaching into the notification tables directly, which nobody knew. Those are exactly the couplings that would have broken the extraction if discovered later.
The service is built to consume events rather than to receive calls. The monolith publishes "order placed", "payment failed" and "account created"; the notification service subscribes and decides what to send. This is deliberate: notifications are a reaction to facts, not a question requiring an answer, so an asynchronous design means the checkout path no longer waits on an email provider and no longer fails when that provider does.
Data moves last. Initially the new service reads notification templates from the existing database. Only once the service has been running in production for several weeks does template ownership move to its own store. Splitting code and data in the same step would have doubled the number of things that could go wrong in one change.
Both paths run in parallel before the switch. The monolith's original code and the new service both process events for a fortnight, with only the monolith actually sending. A comparison job checks that both would have sent the same messages, and it finds three edge cases in undocumented behaviour — a suppression rule for test accounts, a locale fallback, and a rate limit for password resets. Every one of those would have become a customer-visible incident on cutover.
The old code is deleted. This is the step that is most often skipped and the only one that realises the benefit. Until it is gone, the team maintains two implementations and the monolith is no smaller.
The extraction takes roughly three months, most of which is building deployment, tracing, contract testing and on-call practice that did not previously exist. The second extraction takes six weeks. That ratio is normal, and a second extraction that is not substantially faster is the clearest available signal that the boundaries are wrong.
16. Frequently asked questions
How many services should we have?
As few as possible while still allowing teams to deploy independently. A useful upper bound is roughly the number of teams multiplied by a small factor; beyond that, operational burden per team grows faster than the benefit. Organisations frequently report consolidating services after an initial period of over-decomposition, and almost never report regretting having too few.
Should we start with microservices on a new project?
Almost never. At the start you understand the domain least, which is exactly when boundary decisions are most likely to be wrong — and boundaries are the most expensive thing to change afterwards. Build a well-modularised monolith with clear internal seams, and extract services when a specific pressure justifies it. Teams that start distributed usually spend their first year moving boundaries across network calls.
Do we need Kubernetes?
You need something that handles deployment, service discovery, health checking and scaling. Kubernetes does this comprehensively at the cost of substantial operational complexity. For a modest number of services, managed container platforms or even well-orchestrated virtual machines are entirely adequate and considerably simpler. Adopt the heavier tooling when the number of services makes the alternative painful, not in anticipation.
How do we handle reporting across services?
Not by querying each service's database. Publish events to a data platform and build reporting there, which keeps service databases private and gives analysts a model designed for their questions rather than for transaction processing. Cross-service reporting is one of the strongest arguments for taking events seriously from the beginning.
What about shared code between services?
Share libraries for technical concerns — logging, tracing, HTTP clients, authentication helpers. Do not share domain models, because a shared model recreates the coupling you were removing: a change to it requires every service to update. Some duplication of domain concepts across services is correct, since each service's view of a concept is legitimately different.
Is a service mesh worth it?
It moves retries, timeouts, circuit breaking, encryption and traffic policy out of application code into infrastructure, which is valuable at scale and adds a meaningful operational layer to learn and debug. Below roughly ten to fifteen services, well-written libraries usually achieve the same result with far less machinery. Adopt it when maintaining those concerns across many services in several languages becomes the larger problem.
How do we handle authentication between services?
Verify the user at the edge, then propagate a signed token carrying identity and permissions that each service validates independently. Services should not trust a header they cannot verify, and they should not call back to an identity service on every request. Between services, mutual authentication with short-lived certificates is the standard approach and is one of the strongest arguments for a service mesh.
What is the single most important thing to get right?
Boundaries. Everything else can be fixed incrementally — you can add tracing, improve deployment, introduce contract testing at any point. Wrong boundaries mean every feature touches several services forever, and correcting them requires moving code, data and ownership simultaneously. Spend the time on this before writing anything, and validate it by modularising the monolith first.
Key takeaways
- Microservices solve an organisational problem. If teams can already deploy independently, you do not need them.
- Build the operational prerequisites first, inside a monolith, where mistakes are cheap.
- Boundaries follow business capabilities, never technical layers. Validate them with your commit history.
- Each service owns its data exclusively. Duplication via events beats a shared database every time.
- Design for partial failure from day one. Timeouts, idempotency, circuit breakers and graceful degradation.
- Never require a coordinated deployment. The moment you do, you have a distributed monolith.
The best microservices architectures are unremarkable: a modest number of well-bounded services, each owned by one team, communicating mostly through events, deployed independently many times a day. Getting there is less about adopting an architecture than about earning the operational discipline it assumes.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch