Skip to main content
Blog

Fully Automated CI/CD with Jenkins: A Practical Pipeline Guide

Last updated Automation

Jenkins has been declared obsolete roughly once a year for a decade and is still running an enormous share of the world's builds. The reason is unglamorous: it does whatever you need, it runs where you need it, and it has a plugin for the thing nobody else supports. The cost is that it will also let you build something unmaintainable, and many teams have.

This guide covers a Jenkins setup that survives contact with a real team — pipelines as code, sensible agents, artefact handling, environment promotion, secrets, and the operational discipline that keeps the whole thing from becoming the system nobody dares touch.


What you will learn
  • Why pipelines belong in the repository, not in the UI
  • Structuring a pipeline that stays readable as it grows
  • Agents, labels and where builds should actually run
  • Artefacts, versioning and promotion between environments
  • Secrets, credentials and approval gates
  • Keeping a Jenkins installation maintainable over years
In this article
  1. What Jenkins actually is
  2. The single most important decision
  3. Anatomy of a pipeline
  4. Controllers and agents
  5. Where builds should run
  6. Triggering builds
  7. The build stage
  8. Testing stages
  9. Artefacts and versioning
  10. Environment promotion
  11. Deployment strategies
  12. Secrets and credentials
  13. Approval gates and audit
  14. Shared libraries
  15. Speed: where pipelines waste time
  16. Operating Jenkins itself
  17. When Jenkins is the wrong choice
  18. Twelve mistakes
  19. A worked example: a service pipeline end to end
  20. Frequently asked questions

1. What Jenkins actually is

Strip away the plugins and Jenkins is a job scheduler with a web interface. It watches for triggers, allocates work to machines, runs steps, records results, and stores outputs.

That generality is both its strength and its trap. It does not impose a model of what a build should look like, which means it fits any workflow — and means nothing prevents a workflow that should never have existed. The teams with maintainable Jenkins installations are the ones that imposed their own discipline, because Jenkins will not do it for them.

The mental model worth holding: Jenkins orchestrates; it should not implement. Your build logic belongs in scripts in your repository that a developer can run locally. Jenkins decides when to run them, where, and what to do with the results. When build logic lives inside Jenkins, nothing can be tested or reproduced outside it, and that is the root of most Jenkins misery.

2. The single most important decision

Define pipelines as code, in the repository. Not in the web interface, not in freestyle jobs, not in a plugin's configuration screen.

The reasons compound:

  • Version controlled. Pipeline changes are reviewed like code, and you can see what changed and why.
  • Branch-aware. A feature branch can modify its own pipeline without affecting anything else.
  • Reproducible. A pipeline configured through a UI exists only in the controller's disk. Restore that from a backup and you have configuration nobody remembers agreeing to.
  • Reviewable. A change to how software is built and deployed deserves the same scrutiny as a change to the software.

Freestyle jobs — configured entirely through the interface — remain in many installations and should be treated as debt. Every one is a piece of critical infrastructure that exists only as form fields, cannot be diffed, and will be reconstructed from memory the day the controller fails.

Multibranch pipeline configuration takes this further: Jenkins scans a repository, discovers branches containing a pipeline definition, and creates jobs automatically. New branch, new pipeline, no manual job creation. Branch deleted, job removed. This is the arrangement worth reaching, and it removes an entire category of routine administration.

3. Anatomy of a pipeline

A pipeline definition describes stages, what runs in each, and the conditions under which each executes. The shape that works for most services:

StagePurposeTypical durationFails the build?
CheckoutFetch source at a specific commitSecondsYes
Static analysisLint, format, type checksUnder a minuteYes
BuildCompile and packageMinutesYes
Unit testsFast, isolatedMinutesYes
Integration testsAgainst real dependenciesLongerYes
Security scanDependencies and codeMinutesOn high severity
Publish artefactPush to a registryMinutesYes
Deploy to stagingAutomaticMinutesYes
Acceptance testsAgainst stagingLongerYes
Deploy to productionGatedMinutesYes

Two principles govern the ordering. Fail fast: put the quickest checks first so a formatting error is reported in thirty seconds rather than after a twenty-minute test suite. Build once: everything after the build stage operates on the same artefact. Rebuilding per environment means the thing you tested is not the thing you deployed, which defeats the purpose of testing.

4. Controllers and agents

The controller hosts the interface, stores configuration and history, and schedules work. The agents execute it.

Nothing should build on the controller. This is the most consistently violated rule in Jenkins installations and the source of the most preventable outages. Builds consume memory and disk unpredictably; a build that fills the disk takes the controller with it, and the controller holds all your history and configuration. Configure zero executors on the controller and enforce it.

Agents connect to the controller and advertise labels describing what they offer — an operating system, an architecture, an installed toolchain, access to a particular network. Pipeline stages request agents by label rather than by name, so a stage requiring a particular platform gets one without knowing which machine it is.

The label discipline that ages well: describe capabilities, not machines. A label naming a specific host means moving that build requires editing every pipeline that referenced it. A label naming what the build needs means adding a second machine with the same label is all it takes to scale.

5. Where builds should run

ApproachAdvantagesDisadvantages
Static agentsSimple; fast start; caches persistEnvironment drift; idle cost; state accumulates between builds
Container agentsClean every time; environment defined in codeStartup overhead; caching needs deliberate design
Cloud agents on demandElastic; pay for useProvisioning latency; more complex configuration
Container per stageEach stage gets exactly the tools it needsPassing artefacts between stages requires care

Containers are the default worth choosing. The environment is declared alongside the pipeline, every build starts identical, and upgrading a toolchain is a one-line change rather than a maintenance window on ten machines. The class of bug where a build passes on one agent and fails on another simply stops existing.

The cost is that caches do not persist automatically. Dependency downloads, compiler caches and container layers all vanish between builds unless you mount external volumes or use a cache service. Getting this right is usually the difference between a three-minute build and a twelve-minute one, and it is worth the afternoon it takes.

Static agents remain sensible where builds need specialised hardware, where a licensed toolchain cannot be containerised, or where the cache genuinely cannot be externalised. Accept that they drift, and rebuild them from configuration management rather than by hand.

6. Triggering builds

Webhooks from the repository are the correct default. A push triggers a build immediately. Polling — where Jenkins asks the repository whether anything changed — wastes resources and adds latency, and exists mainly because webhooks require the controller to be reachable, which not every network permits.

Scheduled builds serve a specific and underrated purpose: catching failures that are not caused by your changes. A nightly build on the main branch catches a dependency that started failing, an expired certificate, or an external service whose behaviour changed. Without it, the first sign is a developer's build failing for reasons unrelated to their work.

Upstream triggers chain pipelines — a library builds, then services depending on it rebuild. Useful and easy to overuse; a deep chain of triggered builds becomes impossible to reason about, and a failure four levels down is reported against a build nobody was watching.

Manual triggers for deployment and operational tasks, which is covered under gates.

7. The build stage

Three properties matter, and all three are about what happens outside Jenkins.

Reproducible. The same commit produces the same output. This requires pinned dependency versions — a build resolving to whatever is latest is a build that can break without anyone changing anything, and diagnosing that is miserable.

Runnable locally. A developer should be able to run the same build on their machine with one command. If the only way to build is through Jenkins, debugging a build failure means pushing commits and waiting, which is a terrible feedback loop. This follows directly from keeping logic in scripts rather than in pipeline steps.

Producing an identified artefact. Whatever comes out should carry a version, the commit it came from, and when it was built. An artefact you cannot trace to a commit is one you cannot investigate.

8. Testing stages

Split by speed and dependency, and run them accordingly.

Unit tests — fast, no external dependencies, run on every commit and every branch. If they take more than a few minutes, that is a problem to solve rather than accept, because slow feedback changes developer behaviour long before it changes the schedule.

Integration tests — against real databases, queues and services, usually started as containers alongside the build. Slower, and the first place flakiness appears. Run on every commit if they fit the time budget, otherwise on merge to main.

Acceptance tests — against a deployed environment, exercising real workflows. Slowest, most valuable, most fragile. Run after deployment to staging.

Parallelise wherever possible. Test suites split across agents cut wall-clock time roughly linearly, and this is usually the largest single improvement available to a slow pipeline.

Flaky tests are an emergency, not an annoyance. A suite that fails randomly one time in ten trains everyone to re-run rather than investigate, and once that habit forms the suite has stopped providing information. Quarantine flaky tests immediately — out of the blocking path, tracked, with an owner — and fix or delete them. A suite everyone trusts is worth more than a suite with better coverage that nobody believes.

9. Artefacts and versioning

Build once, promote the same artefact. The single most important principle in the whole pipeline, and the one that separates deployment systems that work from ones that produce mysterious environment-specific failures.

Concretely: the build stage produces an artefact and pushes it to a registry. Every subsequent stage — staging deployment, testing, production deployment — references that same artefact by its immutable identifier. Nothing is ever rebuilt for an environment.

This requires that configuration is injected at deployment, not baked in at build. The artefact must be environment-agnostic, receiving its endpoints, credentials and settings from the environment it lands in. An artefact containing a staging URL is an artefact that cannot be promoted.

Versioning that works: an immutable identifier per build, traceable to a commit, never reused. Semantic versions for anything consumed by others, with the build metadata attached. Never a floating tag like "latest" in a deployment — it makes the question "what is running in production" unanswerable.

10. Environment promotion

A typical progression, with the automation level rising as consequences do.

EnvironmentTriggered byGateVerified by
Ephemeral / previewAny branch pushNoneAutomated tests
StagingMerge to mainAutomaticAcceptance tests
ProductionManual or automaticApproval, if requiredSmoke tests, monitoring

Ephemeral environments per branch are the highest-value addition to a pipeline that does not have them. A reviewer can see the change running rather than reading a diff, and product people can try it before merge. They must be genuinely disposable — created on push, destroyed on merge or after a few days idle — or they become a fleet of orphaned environments costing money.

Whether production deployment is automatic or gated is a team decision with a technical prerequisite: automatic deployment requires confidence in the tests and a fast, reliable rollback. Without both, a gate is honest rather than timid.

11. Deployment strategies

The pipeline's final stage should implement a strategy that limits blast radius.

Rolling replaces instances gradually, with health checks between batches. Simple and adequate for most services. Note that two versions run simultaneously during the roll, which the application must tolerate.

Blue-green runs a complete second environment, tests it, and switches traffic. Rollback is a switch back, which makes it the fastest recovery available. It costs double the infrastructure during deployment.

Canary sends a small fraction of traffic to the new version, watches metrics, and proceeds or rolls back. The best risk reduction and the most operational machinery, since it requires meaningful automated comparison of the two populations.

Whatever the strategy, three things must be in place: health checks that actually reflect readiness, not merely that a process started; automated rollback triggered by failed checks rather than by a person noticing; and database migrations decoupled from deployment, applied separately and backwards-compatibly, because a schema change that only works with the new code makes rollback impossible.

12. Secrets and credentials

Jenkins stores credentials and injects them into builds, and how you use that facility determines a large part of your security posture.

Never in the pipeline file. It is in version control, visible to everyone with repository access, and preserved in history forever.

Scope credentials narrowly. Per folder, per job, per environment. A global credential is available to every build, which means any pipeline anyone adds can use your production deployment key.

Prefer short-lived tokens. Where the target system supports issuing temporary credentials, use them. A long-lived key in a credential store is still a long-lived key.

Mask them in output. Jenkins does this for known credentials, and it fails when a credential is transformed — encoded, embedded in a URL, or passed through a script that reformats it. Assume some will leak into logs and rotate accordingly.

External secret stores are better than Jenkins credentials for anything sensitive: rotation, audit and access control belong to a system built for them, and Jenkins fetches at runtime rather than storing.

13. Approval gates and audit

Some deployments need a human decision, whether for regulatory reasons or because the team is not yet ready to automate.

A gate should record who approved, when, and what they approved — the specific artefact and commit. A gate that logs "someone clicked proceed" satisfies nobody's audit requirement.

Two practical cautions. Restrict who can approve, because a gate anyone can click is not a control. And set a timeout, because a pipeline waiting indefinitely holds an executor and clutters the build queue; expiring after a period and requiring a re-run is cleaner.

The uncomfortable truth about gates: they mostly do not catch problems. The person approving rarely has information the automated tests lacked. They are valuable where genuine external coordination is needed — a maintenance window, a customer notification, a regulatory sign-off — and they are theatre when used as a substitute for confidence in testing. Being honest about which situation you are in tends to improve both.

14. Shared libraries

Once several pipelines exist, they repeat themselves. A shared library — pipeline code in its own repository, loaded by pipelines that need it — removes that duplication.

What belongs in one: deployment logic used by many services, notification formatting, artefact publishing conventions, standard security scanning, and any organisational policy that should apply everywhere.

The failure mode is a library that grows into a framework, where every pipeline is a single call to a function taking thirty parameters. It becomes impossible to understand what any given service actually does, and changing the library risks every pipeline simultaneously.

The balance that holds up: libraries provide steps; pipelines compose them. A pipeline should still read as a sequence of recognisable stages, with the library supplying the implementation of each. And version the library, so a pipeline pins a version rather than tracking the tip — otherwise a library change breaks every build at once, which is exactly the failure you built it to avoid.

15. Speed: where pipelines waste time

Slow pipelines change behaviour long before they change delivery dates. Developers batch changes, stop running the full suite locally, and context-switch while waiting. The common causes, roughly in order of impact:

  • No caching. Downloading dependencies on every build, in every stage. Frequently the largest single waste.
  • Serial stages that could be parallel. Lint, unit tests and security scans have no dependency on each other.
  • Rebuilding between stages. Build once, pass the artefact forward.
  • Running everything on every change. A documentation-only commit does not need the full integration suite.
  • Agent provisioning overhead. Warm pools or pre-pulled images.
  • A single monolithic test suite that could be split across agents.

The target worth aiming at: feedback on a commit within ten minutes. Beyond that, developers stop waiting and the pipeline becomes something that reports failures later rather than something that guards the main branch.

16. Operating Jenkins itself

Jenkins is infrastructure and needs the same care as anything else you run.

Back up the configuration directory and test restoring it. It holds every job, credential and setting. Teams discover their backup does not work at the worst possible moment.

Manage plugins deliberately. Plugins are the reason Jenkins can do anything and the reason installations become fragile. Install what you need, remove what you stopped using, and update on a schedule rather than reactively. A controller with two hundred plugins has two hundred sources of incompatibility on the next upgrade.

Configure retention. Build history and artefacts grow without limit. Keep the last N builds, or builds from the last N days, and push artefacts to a proper registry rather than storing them on the controller.

Watch the queue. Builds waiting for agents is the clearest signal of under-capacity, and it is the metric that correlates most directly with developer frustration.

Define the controller as code where possible. Configuration-as-code support lets the installation itself be version controlled and rebuilt, which converts disaster recovery from an archaeology exercise into a deployment.

17. When Jenkins is the wrong choice

Honest assessment, because the alternatives have improved considerably.

Jenkins fits when you need builds on your own infrastructure or in a restricted network; you have complex orchestration across many systems; you need a specific integration that only exists as a Jenkins plugin; or you have significant existing investment that works.

Something else probably fits better when your code is hosted somewhere with integrated pipelines and your needs are ordinary — the integration is tighter and there is no server to operate; when you have nobody to maintain a controller, because an unmaintained Jenkins degrades into a liability; or when you are starting fresh with straightforward requirements.

The deciding question is usually operational rather than technical: does someone own this installation? Jenkins rewards ownership and punishes neglect more than hosted alternatives do.

18. Twelve mistakes

  1. Freestyle jobs. Critical configuration existing only as form fields.
  2. Building on the controller. One runaway build takes down everything.
  3. Build logic inside Jenkins. Nothing testable or reproducible locally.
  4. Rebuilding per environment. You did not deploy what you tested.
  5. Configuration baked into artefacts. Makes promotion impossible.
  6. Global credentials. Every pipeline can use your production key.
  7. Tolerating flaky tests. Trains everyone to ignore failures.
  8. No caching. The most common cause of slow pipelines.
  9. Labels naming machines rather than capabilities.
  10. Unbounded retention. A controller that fills its own disk.
  11. Plugin sprawl. Every upgrade becomes a research project.
  12. Untested backups. Discovered at the worst possible moment.

19. A worked example: a service pipeline end to end

A backend service, containerised, deployed to an orchestrated cluster, with a team of eight.

Structure. The pipeline definition lives in the repository. Multibranch configuration means every branch gets a pipeline automatically. Build logic lives in a handful of scripts in the repository that a developer runs locally with the same commands the pipeline uses — this was the deliberate first decision, and it is why debugging build failures does not require pushing commits.

On every push to any branch. Checkout, then three stages in parallel: linting and formatting, static type checking, and dependency vulnerability scanning. All three complete in under ninety seconds. Then unit tests, split across four agents, four minutes. Then the container image is built and pushed to the registry tagged with the commit identifier. Total: about seven minutes to a published artefact.

For pull requests, additionally. The image is deployed to an ephemeral namespace with a generated URL posted as a comment on the pull request. Integration tests run against it with a real database started alongside. The environment is destroyed when the pull request closes or after three days idle. This addition changed review quality more than anything else the team did — reviewers look at the running change rather than only the diff.

On merge to main. The same image — not rebuilt — is deployed to staging. Acceptance tests run against staging, about eleven minutes. On success, the image is tagged as a release candidate. Nothing beyond this point rebuilds anything.

Production. A gate, currently manual, restricted to four people, recording who approved which artefact and when. Deployment is rolling with health checks, followed by smoke tests against production. If smoke tests fail, the previous version is redeployed automatically. Database migrations run as a separate step before the deployment and are required to be backwards-compatible, which is enforced by a check in the pull request pipeline rather than by convention.

Agents. Container agents on the cluster, provisioned per build. Dependency caches on a persistent volume mounted into build containers — that mount is worth about four minutes per build, which across roughly forty builds a day is a meaningful fraction of a working day recovered.

Secrets. Fetched at runtime from an external store using a short-lived token, scoped so the staging pipeline cannot retrieve production credentials. Jenkins itself holds only the token that lets it authenticate to that store.

What went wrong and what fixed it. Early on, integration tests were flaky at roughly one failure in eight, always in the same three tests, all timing-related. The team's initial response was a retry, which hid the problem and doubled the stage duration. Moving those three tests to a quarantine suite that runs but does not block, with a named owner and a two-week deadline, produced fixes within a week — the deadline mattered more than the mechanism.

Second, the pipeline crept from seven minutes to nineteen over four months as stages were added without anyone watching the total. A single afternoon of parallelising three sequential stages and adding a filter so documentation-only commits skip integration tests brought it back to nine. The lesson the team took was to display total pipeline duration on the same dashboard as the failure rate, because a number nobody sees is a number nobody defends.

Third, an upgrade broke a plugin that a deployment step depended on, and the controller had been upgraded without a staging controller to test against. The installation now runs from configuration-as-code with a second controller used to validate upgrades, which took two days to set up and has prevented three similar incidents since.

20. Frequently asked questions

Should we still use Jenkins in a world of hosted CI?

If you need builds on your own infrastructure, have complex cross-system orchestration, or depend on an integration that only exists as a plugin, yes. If your requirements are ordinary and your code is hosted somewhere with integrated pipelines, the hosted option is usually less work. The deciding factor is whether someone will own the installation — an unmaintained Jenkins becomes a liability faster than most infrastructure.

How do we migrate away from freestyle jobs?

Incrementally, starting with the jobs that change most often. Write the pipeline definition alongside the existing job, run both until the outputs match, then disable the freestyle version. Attempting a big-bang conversion of a hundred jobs invariably stalls, and the jobs that matter are a small fraction of the total.

How long should a pipeline take?

Under ten minutes to feedback on a commit is the target worth defending, because beyond that developers stop waiting and the pipeline stops functioning as a guard on the main branch. Deployment pipelines can be longer since nobody is blocked on them. If your commit pipeline is thirty minutes, caching and parallelisation will usually recover most of it.

Should production deployment be automatic?

It requires two things: tests you genuinely trust and a fast, automated rollback. With both, automatic deployment is safer than manual because it removes the batching that makes each release larger and riskier. Without either, a gate is honest. The useful move is to work toward the prerequisites rather than debating the gate.

Where should build logic live?

In scripts in your repository that developers run locally with the same commands the pipeline uses. Jenkins should call them and handle orchestration, artefacts and reporting. When logic lives inside pipeline steps or plugin configuration, nothing can be reproduced or debugged outside Jenkins, which turns every build failure into a push-and-wait cycle.

How do we handle a monorepo?

Detect which paths changed and run only the affected pipelines. Building everything on every commit does not scale past a certain size, and the detection logic is usually simpler than teams expect — a mapping from directories to pipelines, plus a rule that shared code triggers everything downstream of it.

What about plugin updates?

On a schedule, tested against a second controller before applying to the one people depend on. Reactive updating — only when something breaks or a vulnerability is announced — means every update is a large jump with a high chance of incompatibility. Keeping the plugin list short is the other half of the answer.

What is the highest-value improvement to an existing pipeline?

Usually dependency caching, which is frequently several minutes per build for an afternoon of work. After that, running independent stages in parallel, and ephemeral environments per pull request — the last one changes review quality rather than speed, which is a different kind of value but often the larger one.

Key takeaways

  • Pipelines as code, in the repository. Everything else follows from this.
  • Jenkins orchestrates; scripts implement. Build logic must run locally.
  • Build once, promote the artefact. Configuration is injected, never baked in.
  • Never build on the controller, and label agents by capability.
  • Flaky tests are an emergency. A suite nobody trusts provides nothing.
  • Under ten minutes to feedback, or developers stop treating it as a guard.

A good Jenkins setup is unremarkable: commits get checked quickly, artefacts are traceable, deployments are boring, and nobody thinks about the controller. Getting there is mostly discipline rather than cleverness — putting the pipeline in the repository, keeping logic outside Jenkins, and treating the installation as infrastructure somebody owns.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch