Skip to main content
Blog

Docker and Kubernetes: Container Orchestration for Production

Last updated Containers

Containers solved a real problem: software that runs on one machine and not another. Orchestration solved a different one: running many containers across many machines without a human deciding where each goes. The two are frequently discussed as one topic, which is unfortunate, because most teams need the first long before they need the second — and adopting the second prematurely is the most common expensive mistake in modern infrastructure.

This guide covers both properly: how images actually work and how to build good ones, what an orchestrator does and the model it imposes, what production genuinely requires beyond a running pod, and how to decide whether you need any of it.


What you will learn
  • What a container actually is, and why images behave the way they do
  • How to build images that are small, fast and safe
  • The orchestration model — declarative state and controllers
  • The objects that matter and what each is for
  • Health, scaling, configuration, storage and networking in practice
  • What production requires, and whether you need an orchestrator at all
In this article
  1. What a container is
  2. Images and layers
  3. Building good images
  4. Why orchestration exists
  5. The declarative model
  6. The objects that matter
  7. Health and lifecycle
  8. Scaling
  9. Configuration and secrets
  10. Networking
  11. Storage and stateful workloads
  12. Security
  13. Observability
  14. What production actually requires
  15. Do you need an orchestrator?
  16. Twelve mistakes
  17. A worked example: one service to production
  18. Frequently asked questions

1. What a container is

A container is a process running on the host kernel with a restricted view of the system. It is not a virtual machine — there is no guest operating system and no hypervisor. The isolation comes from kernel features that limit what the process can see and how much it can consume.

Two mechanisms do the work. Namespaces control visibility: the process sees its own filesystem, its own network interfaces, its own process tree. Control groups limit consumption: how much processor time and memory it may use.

Three consequences follow directly, and they explain most container behaviour:

  • Startup is fast — starting a process rather than booting an operating system.
  • Isolation is weaker than a virtual machine. A kernel vulnerability affects everything on the host, which matters when running untrusted code.
  • The kernel is shared, so containers must be compatible with the host's kernel. This is why Linux containers require a Linux kernel, provided by a lightweight virtual machine on other operating systems.

2. Images and layers

An image is a stack of read-only filesystem layers plus metadata describing how to run it. Each build instruction that changes the filesystem produces a layer, and layers are shared between images — pulling a second image based on the same foundation downloads only the differences.

Two properties of layers explain most build behaviour and most build problems.

Caching is sequential. A layer is reused only if every layer before it is unchanged. This is why instruction order matters enormously: put the things that change rarely — installing dependencies — before the things that change constantly — copying source code. Reversing that order means reinstalling dependencies on every build.

Deletion does not reduce size. Removing a file in a later layer hides it; the data remains in the earlier layer and in the image. Adding a secret and deleting it in a subsequent instruction leaves the secret in the image, retrievable by anyone who pulls it. This surprises people regularly and has leaked real credentials.

The other property worth internalising is that the container filesystem is ephemeral. Anything written inside a container is lost when it is replaced, which is the correct default and the reason state must live somewhere designed for it.

3. Building good images

A well-built image is small, builds quickly, and contains nothing that is not needed to run.

  • Use multi-stage builds. Compile in a stage containing your toolchain, then copy only the artefact into a minimal runtime stage. This routinely reduces image size by an order of magnitude and removes compilers and build tools from production.
  • Order instructions by change frequency. Dependency manifests copied and installed before application source, so a code change does not invalidate the dependency layer.
  • Choose a minimal base. Slim or distroless variants contain fewer packages, which means a smaller attack surface and fewer vulnerabilities reported by scanners. The trade-off is that debugging inside a container without a shell requires different techniques.
  • Pin versions. A base image tag that moves means builds are not reproducible and a rebuild can change behaviour without a code change.
  • Run as a non-root user. A process running as root inside a container is a privilege escalation waiting for a kernel vulnerability.
  • Exclude build context you do not need. An ignore file prevents sending large or sensitive directories to the build, which speeds builds and prevents accidental inclusion.
  • Never place secrets in an image. Use build-time secret mechanisms or inject at runtime. Layers preserve everything.
  • Handle signals properly. A process that ignores termination signals will be killed forcibly after a grace period, cutting off in-flight requests.

4. Why orchestration exists

Running one container is straightforward. The problems appear at scale, and each orchestration feature exists to solve one of them.

ProblemWhat the orchestrator does
Which machine should this run on?Schedules based on available resources and constraints
A container crashedRestarts it automatically
A machine failedReschedules its workloads elsewhere
How do services find each other?Provides stable names and load balancing
How do we deploy without downtime?Rolling updates with health verification
Traffic increasedScales replicas on measured demand
Configuration differs per environmentInjects configuration and secrets separately from images
A deployment went wrongRolls back to the previous known-good version

Each of these can be solved individually with simpler tools. The argument for an orchestrator is that solving all of them consistently, across many services, is what it does — and the argument against is that it is a substantial system to learn and operate, which is not free.

5. The declarative model

The central idea is worth stating precisely, because it explains behaviour that otherwise seems mysterious.

You do not tell the system what to do. You describe the state you want — three replicas of this image, exposed on this port — and controllers continuously compare actual state against desired state and act to close the gap. This runs forever, not once.

Three implications follow:

  • Manual changes are reverted. Deleting a pod belonging to a deployment creates a replacement, because desired state still says three. To stop something, change the declaration.
  • Convergence is eventual and can fail. Requesting more resources than the cluster has produces a pod that stays pending indefinitely rather than an error. The system is not failing; it is waiting.
  • The declaration is the source of truth, which is why storing it in version control and applying it through a pipeline is the natural operating model. Changes made directly against a cluster are undocumented and will eventually be overwritten.

6. The objects that matter

The object catalogue is large; a small subset covers most work.

ObjectPurpose
PodOne or more containers scheduled together sharing a network address. The unit of scheduling, rarely created directly.
DeploymentMaintains a number of identical pods and manages rolling updates and rollback. The workhorse for stateless services.
ServiceA stable network name and load balancer for a changing set of pods.
Ingress or GatewayRoutes external HTTP traffic to services, handling hostnames, paths and certificates.
ConfigMapNon-sensitive configuration injected as environment variables or files.
SecretSensitive configuration, with different access controls. Encoded by default, not encrypted.
NamespaceA grouping boundary for names, quotas and access control.
StatefulSetFor workloads needing stable identity and persistent storage per instance.
Job and CronJobRun-to-completion work, once or on a schedule.
PersistentVolumeClaimA request for durable storage, satisfied by the cluster's storage provider.

The relationship worth understanding first: a deployment creates and manages pods, a service gives them a stable address, and an ingress routes outside traffic to that service. That chain covers the majority of what a typical web service needs.

7. Health and lifecycle

Health checking is where most production problems originate, because the defaults are permissive and the distinctions are subtle.

Three kinds of check, each answering a different question:

  • Startup: has the application finished starting? Protects slow-starting applications from being killed before they are ready.
  • Readiness: can this instance serve traffic right now? Failing removes it from the load balancer without restarting it.
  • Liveness: is this instance unrecoverably broken? Failing restarts the container.

The distinction between readiness and liveness matters more than any other configuration detail. A readiness check that includes a database dependency is correct — an instance that cannot reach the database should not receive traffic. A liveness check that includes the same dependency is dangerous: when the database has a brief problem, every instance restarts simultaneously, turning a short degradation into an outage.

The rule: liveness checks should test only whether this process is broken, not whether its dependencies are healthy.

Shutdown deserves equal attention. When a pod is terminated, it is removed from the load balancer and sent a termination signal, then forcibly killed after a grace period. An application that exits immediately on that signal drops in-flight requests. The correct behaviour is to stop accepting new work, finish what is in progress, and then exit — with a grace period longer than the longest expected request.

8. Scaling

Three kinds, frequently confused.

Horizontal pod scaling adds or removes replicas based on a metric. Processor utilisation is the default and is frequently the wrong signal — a service bounded by external calls will not show processor pressure while queueing badly. Scale on request rate, queue depth or latency where you can.

Vertical scaling adjusts the resources allocated to each pod. Useful for right-sizing, and it typically requires restarting the pod, which limits its usefulness for reactive scaling.

Cluster scaling adds or removes machines when pods cannot be scheduled. Necessary for horizontal scaling to actually work beyond current capacity, and it introduces a delay — new machines take minutes to become available, which means bursty traffic needs headroom rather than reactive scaling alone.

Underneath all three sits the most consequential configuration in the system: resource requests and limits. Requests are what the scheduler uses to decide placement; limits are enforced ceilings. Setting requests too low causes overcommitment and unpredictable performance. Omitting them entirely means the scheduler is guessing, which is the most common cause of mysterious instability in a cluster. Memory limits are enforced by termination, so a limit set below actual usage produces repeated restarts that look like an application bug.

9. Configuration and secrets

The rule is that images are identical across environments and configuration is injected. This is what makes "build once, promote the same artefact" possible.

Configuration objects can be exposed as environment variables or mounted as files. Files are preferable for anything large, anything structured, and anything that should be updatable without restarting the pod.

Secrets deserve care, because the built-in mechanism is weaker than its name suggests: values are encoded rather than encrypted by default, and anyone who can read the object can read the secret. Production practice requires enabling encryption at rest, restricting access properly, and — better — integrating an external secrets manager so the cluster holds a reference rather than the value.

Two operational details that cause confusion: a configuration change does not automatically restart pods, so a deployment must be triggered for the new values to take effect unless the application watches for changes; and secrets mounted as files do update in place, which applications can take advantage of if they re-read them.

10. Networking

The networking model has a few rules that, once understood, make the rest predictable.

Every pod gets its own address and can reach every other pod directly. There is no network address translation between pods. Pod addresses are ephemeral — they change on every restart — which is exactly why services exist.

A service provides a stable name and address, load balancing across the pods matching its selector. Internal name resolution means one service can reach another by name without knowing anything about placement.

External traffic arrives through an ingress or gateway, which handles hostname and path routing, certificates and often rate limiting. A load balancer service type exposes a single service directly and is more expensive at scale, since it provisions infrastructure per service.

Network policies restrict which pods may talk to which. The default is that everything can reach everything, which means a compromised component has the whole cluster available to it. Defining policies — at minimum, restricting access to data stores — is one of the highest-value security measures available and is frequently skipped.

A service mesh adds mutual authentication, retries, timeouts, circuit breaking and detailed traffic metrics at the infrastructure layer. Genuinely valuable at scale and a substantial operational addition; below roughly ten to fifteen services, well-written libraries achieve the same with far less machinery.

11. Storage and stateful workloads

Containers are ephemeral; persistent data needs storage that outlives them. A claim requests durable storage, and the cluster's provider satisfies it.

The properties that constrain design: most block storage can be attached to only one node at a time, which means pods using it cannot be spread across nodes freely; shared filesystems allow multiple readers and writers at higher cost and lower performance; and storage is usually zone-bound, so a pod using it must be scheduled in the same zone.

For workloads with genuine identity — database replicas, message brokers — a stateful set provides stable names, ordered startup and per-instance storage that survives rescheduling.

The honest guidance most teams should hear: use managed data services rather than running databases in your cluster, unless you have a specific reason and the expertise to operate them. Backups, failover, upgrades and performance tuning for a database are a specialist discipline, and the orchestrator does not make them easier. Operators automate much of this and are excellent when you genuinely need self-hosting, but they do not remove the need to understand the database.

12. Security

  • Run as non-root, with a read-only root filesystem and all capabilities dropped unless specifically required. These settings prevent a large class of container escape.
  • Enforce policy at admission, so non-compliant workloads are rejected rather than reported. Preventive beats detective.
  • Scan images in the pipeline and continuously afterwards, since vulnerabilities are discovered in images already deployed.
  • Use workload identity rather than static credentials, so pods obtain short-lived tokens for the resources they need.
  • Restrict permissions properly. The default service account should have no meaningful access, and each workload should have exactly what it requires.
  • Apply network policies, starting with default-deny in sensitive namespaces.
  • Protect the control plane. Access to it is access to everything; treat it as your most sensitive endpoint.
  • Keep versions current. Clusters accumulate known vulnerabilities quickly, and upgrades become harder the longer they are deferred.

13. Observability

Containers are replaced constantly, which changes how observability must work: nothing useful can live on the instance.

Logs should go to standard output and be collected by an agent, never written to files inside the container. Structured output with the service, version and a correlation identifier on every line is what makes them queryable.

Metrics exposed by each workload and scraped centrally. The four that answer most questions are request rate, error rate, latency distribution and saturation — and they should be standardised so every service reports them identically.

Traces across service boundaries, which is the only practical way to understand a slow request crossing several components.

Cluster events are the most under-used signal. When a pod is not running, the events explain why — insufficient resources, image pull failure, failing health check — and reading them is usually faster than any other diagnostic step.

Alert on user-visible symptoms rather than on pod health. Pods restarting is normal; a journey failing is not. Alerting per pod produces noise that teaches everyone to ignore alerts.

14. What production actually requires

A workload running is not a workload in production. The gap is consistent enough to enumerate.

  1. Resource requests and limits set from measured usage, not guessed.
  2. Readiness and liveness checks, correctly distinguished.
  3. Graceful shutdown handling with an adequate grace period.
  4. Multiple replicas spread across zones, with a disruption budget so maintenance cannot take them all down at once.
  5. Rolling updates configured so a broken version cannot replace a working one wholesale.
  6. A rollback path that has been exercised.
  7. Configuration and secrets injected, never baked into images.
  8. Network policies restricting access to data stores.
  9. Logs, metrics and traces flowing to a central place.
  10. Alerts on user-visible symptoms with a named owner.
  11. Image scanning in the pipeline and continuously afterwards.
  12. Cluster and node upgrades planned rather than deferred indefinitely.

15. Do you need an orchestrator?

The honest answer for many teams is no, and the question deserves asking before the commitment rather than after.

You probably do not need one if you run a handful of services, deploy a few times a week, have no dedicated infrastructure capability, and have no requirement for portability across providers. Managed container platforms handle deployment, scaling and health checking with a fraction of the concepts to learn.

You probably do need one if you run dozens of services across multiple teams, need sophisticated deployment patterns, have genuine multi-cloud or on-premises requirements, already have the expertise, or depend on the ecosystem of tooling built around it.

The middle ground is large and is where most regret lives. Adopting an orchestrator for five services costs a great deal of learning and operational burden for benefits you could obtain more simply. Migrating to one later is entirely feasible — containers are portable, which is the point — so deferring the decision costs less than making it prematurely.

16. Twelve mistakes

  1. No resource requests or limits. The scheduler guesses; performance becomes unpredictable.
  2. Liveness checks that test dependencies. A brief database problem restarts every instance simultaneously.
  3. Ignoring termination signals. In-flight requests dropped on every deployment.
  4. Secrets baked into images. Layers preserve them permanently.
  5. Running as root. A privilege escalation waiting for a kernel vulnerability.
  6. Moving image tags. Nobody can say what is actually deployed.
  7. Default-allow networking. A compromised component reaches everything.
  8. Writing state to the container filesystem. Lost on every replacement.
  9. Databases in the cluster without expertise. Backups, failover and upgrades are a specialist discipline.
  10. Changes applied directly to the cluster. Undocumented, unreviewed, eventually overwritten.
  11. Alerting on pod restarts. Noise that trains everyone to ignore alerts.
  12. Adopting an orchestrator for five services. Substantial operational weight for no benefit.

17. A worked example: one service to production

Consider taking a straightforward web service from a repository to a production deployment, with the decisions made deliberately.

The image comes first, and it is multi-stage. The build stage contains the compiler and dependencies; the runtime stage contains the binary, a minimal base and nothing else. Dependency installation precedes source copying, so an ordinary code change rebuilds in seconds rather than minutes. The image runs as a non-root user, and the resulting artefact is a fraction of the size the naive single-stage version produced.

Health endpoints are added to the application, not configured around it. A readiness endpoint checks that the service can reach its database, because an instance that cannot should not receive traffic. A liveness endpoint checks only that the process is responsive, deliberately excluding dependencies — the distinction that prevents a brief database problem from cascading into a full restart storm.

Graceful shutdown is implemented before the first deployment. On receiving a termination signal, the service stops accepting new connections, allows in-flight requests to complete, and exits. The grace period is set comfortably above the longest expected request. Without this, every deployment drops a small number of requests, which is easy to miss and difficult to diagnose later.

Resources are set from measurement, not estimation. The service runs under realistic load in a staging environment for a few days; requests are set near observed steady usage and limits above observed peak. Memory limits in particular are set with headroom, because exceeding a memory limit terminates the container and the resulting restart loop looks exactly like an application bug.

The deployment declares three replicas spread across zones, with a disruption budget ensuring at least two remain available during node maintenance. The rolling update configuration allows one pod to be unavailable and one surge pod, so a broken image cannot replace the working version wholesale — the rollout stops when the new pods fail readiness.

Configuration is injected, secrets come from an external manager, and the same image that ran in staging is promoted unchanged to production. A network policy restricts the service to reaching only its database and the services it genuinely calls, which takes an hour and eliminates lateral movement as a concern.

Observability is wired before launch, not after. Logs to standard output with a correlation identifier, the four standard metrics exposed and scraped, traces across the database boundary, and alerts on request error rate and latency rather than on pod restarts.

The rollback is exercised deliberately before it is needed — a previous version deployed, verified and rolled forward again — so that the procedure is familiar rather than improvised during an incident.

None of this is exotic, and the whole sequence is perhaps two days of work. The distance between a service that runs and a service that is genuinely in production is almost entirely made up of these unglamorous items.

18. Frequently asked questions

Do we need Kubernetes?

Only if the number of services, the sophistication of your deployment needs, or a genuine portability requirement justifies the operational weight. Below roughly ten to fifteen services, managed container platforms deliver deployment, scaling and health checking with far fewer concepts. Containers remain portable either way, so starting simpler does not lock you out of adopting an orchestrator later.

Managed or self-hosted cluster?

Managed, in almost every case. Operating a control plane — availability, upgrades, certificates, etcd backups — is specialist work with no differentiating value. Self-host only when regulation, air-gapped requirements or genuinely unusual constraints demand it, and budget for the ongoing operational commitment rather than treating it as a setup cost.

Should we run our database in the cluster?

Usually not. Managed database services handle backups, failover, patching and performance tuning, all of which are specialist disciplines the orchestrator does not simplify. Operators make self-hosting viable and are appropriate when you have a specific requirement and the expertise to operate the database itself. The orchestrator manages containers well; it does not manage a database for you.

How do we manage configuration across environments?

A templating or overlay tool that keeps a shared base with per-environment differences, stored in version control and applied through a pipeline. Avoid maintaining fully separate manifests per environment, which diverge silently. Whichever tool you use, the important property is that the cluster's state is derived from the repository rather than from anyone's terminal history.

What causes most production incidents?

Resource limits set incorrectly, liveness checks that test dependencies, and applications that do not shut down gracefully. All three are configuration rather than platform problems, all three are avoidable, and all three produce symptoms that look like something else — which is why they consume disproportionate diagnostic time.

Is a service mesh worth adopting?

At scale, yes: mutual authentication, retries, timeouts and detailed traffic metrics at the infrastructure layer, applied consistently across services in any language. Below roughly ten to fifteen services, well-written libraries achieve the same result with substantially less machinery to learn and debug. Adopt it when maintaining those concerns across many services becomes the larger problem.

How do we keep clusters up to date?

Treat upgrades as routine scheduled work rather than as projects. Versions are supported for limited periods, and deferring upgrades makes each one harder because you eventually cross multiple versions at once. Upgrade non-production first, watch for deprecated interfaces in your manifests, and keep node images current — an out-of-date node pool accumulates vulnerabilities as quickly as the control plane does.

What is the highest-value thing to fix in an existing cluster?

Resource requests and limits derived from measurement, and correctly distinguished health checks. Together they account for the majority of unexplained instability, and both are a day of work. Network policies restricting access to data stores are the strongest security improvement for comparable effort.

Key takeaways

  • Containers and orchestration are separate decisions. Most teams need the first long before the second.
  • Layers explain build behaviour. Order by change frequency, and never let a secret enter a layer.
  • Declarative means continuous reconciliation. Manual changes are reverted; the declaration is the truth.
  • Liveness must not test dependencies. The single most damaging misconfiguration available.
  • Requests and limits from measurement. Guessed values are the leading cause of mysterious instability.
  • Default-deny networking and non-root containers are the highest-value security settings for the effort.

Orchestration is a powerful tool that solves problems you may not have. Containers are worth adopting almost universally; the platform above them is worth adopting when the number of services makes the alternative genuinely painful — and not a service earlier.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch