Skip to main content
Blog

SonarQube: Continuous Code Quality and Security in Your Pipeline

Last updated Code Quality

Static analysis tools have a reputation problem, and they have earned it. A team installs one, points it at an existing codebase, receives fourteen thousand findings, argues about a few of them, marks most as won't-fix, and quietly stops looking. The tool was not wrong; the adoption was. The single decision that separates static analysis that works from static analysis that gets ignored is whether you gate on the entire codebase or only on the code you just wrote.

This guide covers using SonarQube properly: what it actually measures, why the clean-as-you-code approach resolves the adoption problem, how to configure quality gates that people respect, and how to interpret the metrics without turning them into targets that get gamed.


What you will learn
  • What static analysis can and cannot detect
  • The issue taxonomy, and which categories deserve which response
  • Why gating on new code is the decision that makes adoption succeed
  • How coverage, duplication and complexity should actually be used
  • Pipeline integration, branch analysis and pull request decoration
  • Tuning rules so findings are trusted rather than dismissed
In this article
  1. What the tool actually does
  2. The limits of static analysis
  3. The issue taxonomy
  4. Security findings and hotspots
  5. Clean as you code
  6. Quality gates
  7. Coverage, honestly
  8. Duplication
  9. Complexity
  10. Technical debt as a measure
  11. Pipeline integration
  12. Tuning the rule set
  13. Adopting it on an existing codebase
  14. Governance across many projects
  15. Twelve mistakes
  16. A worked example: adoption over one quarter
  17. Frequently asked questions

1. What the tool actually does

SonarQube analyses source code without executing it, applying a rule set per language, and aggregates the results into a picture of a project's state over time.

Concretely, it parses the code into a syntax tree, walks that tree applying rules, and performs data-flow analysis to track how values move — which is how it detects that a variable may be null on one path, or that user input reaches a database query without sanitisation. It also ingests test coverage reports produced by your own test tooling; it does not measure coverage itself.

What it produces: issues categorised by type and severity, security findings, duplication measurements, complexity metrics, and an estimate of remediation effort. The important part is that all of these can be scoped to new code — code added or changed since a defined baseline — which is what makes the tool usable on a codebase with history.

2. The limits of static analysis

Knowing what it cannot see prevents both over-reliance and unfair dismissal.

It finds patterns known to be problematic: null dereferences on reachable paths, resources not closed, injection risks where tainted input reaches a sink, unreachable code, misuse of language and library idioms, and duplicated blocks.

It does not find logic errors where the code correctly implements the wrong requirement, architectural problems, performance issues that depend on data volume, race conditions in most cases, business rule violations, or anything requiring knowledge of intent.

Two consequences follow. It is a complement to tests and review, never a replacement — a codebase with a perfect quality gate and no tests is not a quality codebase. And false positives are inherent: a tool reasoning about all possible paths without executing them will sometimes flag a path that cannot occur in practice. That is not a defect to be eliminated but a cost to be managed through tuning.

3. The issue taxonomy

TypeMeansResponse
BugCode that will demonstrably behave incorrectlyFix. These are the highest-value findings.
VulnerabilityA security weakness with a plausible exploit pathFix, prioritised by exploitability
Security hotspotCode needing human judgement about whether it is safeReview and mark, do not blindly fix
Code smellMaintainability problem — confusing, duplicated, over-complexFix in new code; batch in old code

The distinction between a vulnerability and a hotspot is the one most often missed, and it is well designed. A vulnerability is something the tool is confident about. A hotspot is a place where security-sensitive code exists and only a human can say whether it is a problem — a disabled certificate check might be a serious flaw or a deliberate choice in a local test harness. Hotspots require review and a recorded decision, which is exactly right, and treating them as issues to be silenced defeats their purpose.

Severity ratings are useful as a rough ordering and should not be treated as absolute. A blocker in code nobody executes matters less than a major issue in an authentication path. Judgement remains necessary.

4. Security findings and hotspots

The security analysis is the part that most justifies the tool for many teams, and it works through taint tracking: following untrusted input from where it enters the application to where it is used in a sensitive operation.

That is how it detects injection risks, path traversal, unvalidated redirects and similar issues — not by pattern matching on function names, but by establishing that a value the user controls reaches a place where control matters. The analysis is genuinely useful and it is bounded: it cannot follow taint through a database round trip, through a message queue, or through dynamic dispatch it cannot resolve.

The practical guidance for security findings: fix vulnerabilities, treating exploitability rather than nominal severity as the ordering; review every hotspot and record the decision, because the recorded reasoning is what makes the next review fast; and remember that this covers code you wrote, not your dependencies — dependency vulnerabilities need separate scanning, and they are frequently the more common real attack path.

5. Clean as you code

This is the central idea, and it is worth stating as a principle rather than a feature.

Attempting to fix an existing codebase's entire backlog is a project that never finishes and never gets funded. Instead, define a baseline — a date, or a version — and apply strict standards only to code written after it. Existing code is measured but not gated.

The consequences are what make it work:

  • Standards are achievable. A developer is asked to clean up the fifty lines they just wrote, not the fifty thousand they inherited.
  • Quality improves where change happens. Code being modified is code that matters; untouched code is not causing problems today.
  • The gate can be strict without being unreasonable, because it applies to a small, fresh, well-understood surface.
  • Improvement compounds automatically. Over a year, the actively developed portion of the codebase becomes clean without anyone running a remediation project.

The contrast is stark. Gating on overall metrics means a change to a legacy file fails the build for problems the author did not create, which produces resentment, exemptions, and eventually a disabled check. Gating on new code produces a standard people accept because it is fair.

6. Quality gates

A quality gate is a set of conditions a project must satisfy to pass. It is the enforcement mechanism, and its design determines whether the tool changes behaviour or generates reports.

A gate that works in practice, all conditions applied to new code only:

ConditionThresholdReasoning
New bugsNoneDemonstrably incorrect code should not merge
New vulnerabilitiesNoneSame, for security
New hotspots reviewedAllA decision recorded, not necessarily a change made
Coverage on new codeA realistic figure for your contextNew code should be tested; the number is negotiable
Duplication on new codeA low percentageCatches copy-paste before it spreads
Maintainability rating on new codeThe top ratingKeeps new code clean without touching old

Two design rules matter more than the specific numbers. The gate must actually block. A gate that reports a failure while allowing the merge is a metric, not a control, and it will be ignored within a month. And the conditions must be few enough to be remembered — a gate with fifteen conditions produces failures nobody can interpret quickly, which creates pressure to bypass it.

7. Coverage, honestly

Test coverage measures which lines executed during the test run. It does not measure whether anything meaningful was asserted, which is why it is simultaneously useful and dangerous.

Useful because it identifies code with no tests at all, which is genuinely informative. Coverage on new code is a reasonable gate condition, because writing tests alongside new code is a habit worth enforcing.

Dangerous as a target, because it is trivially gamed. A test that calls every method and asserts nothing produces full coverage and zero value. Teams under coverage pressure reliably produce such tests, and the resulting suite is worse than no suite because it creates false confidence and slows the pipeline.

Practical guidance: gate on coverage of new code at a figure your team genuinely accepts; never set an overall target for an existing codebase; ignore small fluctuations; and look at branch coverage rather than line coverage, since a line executed in one direction of a condition is only half tested. Most importantly, treat low coverage as a question to investigate rather than a number to raise.

8. Duplication

Duplication detection finds repeated blocks across the codebase. It is one of the more reliable signals, because duplication has a specific cost: a change must be made in several places, and eventually one is missed.

The nuance is that not all duplication is bad. Two pieces of code that look similar today and will diverge tomorrow are correctly separate — premature unification couples things that need to change independently. The traditional guidance to wait for a third occurrence before extracting is sound.

Where duplication genuinely signals a problem: the same business rule implemented in several places, near-identical files created by copying, and repeated blocks within a single file that indicate a missing abstraction. Gating on duplication in new code catches copy-paste at the moment it happens, which is when it is cheapest to address.

9. Complexity

Two measures, and the difference is worth understanding.

Cyclomatic complexity counts independent paths through a function, which corresponds roughly to the number of tests needed to cover it. It is precise and it treats a long switch statement as highly complex even though such code is easy to read.

Cognitive complexity attempts to measure how hard code is for a human to understand, penalising nesting more heavily than sequence and treating shorthand structures more kindly. It correlates better with the intuition of "this function is difficult", which makes it the more useful of the two for review purposes.

High complexity is a signal rather than a defect. A function with deeply nested conditions is harder to test, harder to change safely, and more likely to contain a bug — but the fix is a design change, not a mechanical split. Use complexity to identify candidates for attention, and apply judgement about whether extraction genuinely improves clarity or merely relocates it.

10. Technical debt as a measure

The tool estimates remediation effort by assigning a nominal time to each issue and summing them, producing a debt figure and a ratio against the estimated cost of writing the code.

Used carefully, this is a reasonable communication device: a ratio that is rising is a signal worth discussing, and a comparison between modules highlights where change is likely to be expensive.

Used carelessly, it produces bad decisions. The per-issue estimates are generic rather than contextual, so the total is indicative at best. And a debt figure presented to leadership as a number to be reduced invites remediation projects that fix cheap issues in untouched code — activity that improves the metric and nothing else.

The honest framing: technical debt in the sense that matters is the gap between how the system is structured and how it now needs to be structured, which no static analyser can measure. The tool measures rule violations, which correlate loosely. Treat the number as a conversation starter rather than a target.

11. Pipeline integration

The tool only changes behaviour if it runs automatically and its result is visible where decisions are made.

The arrangement that works:

  • Analysis on every pull request, comparing against the target branch so the developer sees only what their change introduced.
  • Results decorated onto the pull request as comments and a status check, rather than living in a separate dashboard nobody opens.
  • The status check blocks merging when the gate fails. This is the step that converts a report into a control.
  • Analysis on the main branch after merge, which updates the project's baseline and history.
  • Coverage reports generated before analysis and passed in, since the tool ingests rather than measures them. A missing coverage report silently produces zero coverage, which is a common and confusing misconfiguration.

Keep the analysis fast enough not to slow the pipeline meaningfully. For large repositories, analysing only changed modules on pull requests and the whole project on the main branch is a reasonable compromise.

12. Tuning the rule set

The default rule set is a starting point and should not survive contact with your codebase unchanged. Untuned rules produce noise, noise produces dismissal, and dismissal makes the whole exercise worthless.

The tuning process that works:

  1. Run against your codebase and read the most frequent rules. A rule producing thousands of findings is either revealing something systemic or does not fit your conventions.
  2. Disable rules that conflict with a deliberate team convention. Arguing with the tool every time is worse than turning it off.
  3. Adjust severities to your context. The defaults are generic; a rule that matters enormously in your domain can be raised, and one that does not can be lowered.
  4. Exclude generated code, vendored dependencies and migrations. Analysing code nobody writes produces findings nobody can act on.
  5. Review the false positive rate periodically. If developers routinely mark findings as false positives, that rule needs attention rather than repeated dismissal.

Aim for a rule set where a finding is credible by default. Developers should approach an issue expecting it to be real; the moment the expectation flips, the tool has stopped working regardless of what it detects.

13. Adopting it on an existing codebase

The sequence that avoids the fourteen-thousand-findings failure:

Analyse without gating first. Get the baseline picture, and resist the urge to act on it. The purpose is information, not a work queue.

Set the baseline to today. Everything before is measured; everything after is gated. This single configuration decision is what makes adoption viable.

Tune the rules against the real findings, as described above, before turning on any gate.

Enable a minimal gate covering new bugs and new vulnerabilities only. Nothing about coverage or maintainability yet. Let the team experience a gate that fails only for things they agree are problems.

Add conditions gradually as the team's confidence grows — hotspot review, then duplication, then coverage on new code, then maintainability. Each addition should be discussed rather than imposed.

Address the historical backlog opportunistically. When a file is being modified anyway, fix its issues. This is far more effective than a remediation project, because it concentrates effort on code that is actually changing.

14. Governance across many projects

With twenty repositories, per-project configuration diverges immediately. Three mechanisms keep it coherent.

Shared quality profiles per language, defined centrally and inherited by projects. A rule change is made once. Projects may extend but should not silently weaken.

Shared quality gates, similarly. A small number — perhaps a standard gate and a stricter one for critical services — rather than one per team.

Portfolio-level reporting so leadership sees the aggregate picture without needing per-project detail. This is genuinely useful for spotting a project drifting, and it is also where the metric-as-target risk is highest, so present trends rather than league tables.

The governance failure to avoid is imposing a gate that teams cannot meet and then granting exemptions. Exemptions accumulate, the standard becomes nominal, and nobody trusts the reporting. A slightly weaker gate that everybody genuinely meets is worth considerably more.

15. Twelve mistakes

  1. Gating on the whole codebase. Unachievable, resented, and eventually disabled.
  2. Not tuning the rule set. Noise trains everyone to dismiss findings.
  3. A gate that reports but does not block. A metric, not a control.
  4. Coverage as a target. Produces assertion-free tests and false confidence.
  5. Missing coverage report in the pipeline. Silently reports zero and blocks everything.
  6. Treating hotspots as issues to silence. Defeats the purpose of the category.
  7. Analysing generated code and dependencies. Findings nobody can act on.
  8. Debt figures presented as a target. Invites remediation of cheap issues in untouched code.
  9. Fifteen gate conditions. Failures nobody can interpret quickly.
  10. Results only in a separate dashboard. Not where the decision is made, therefore not seen.
  11. Exemptions rather than a realistic gate. The standard becomes nominal.
  12. Treating it as a substitute for review and tests. It detects patterns, not wrongness.

16. A worked example: adoption over one quarter

Consider a team of twelve with a six-year-old codebase, no static analysis, and a general sense that quality is declining. Here is a sequence that works, deliberately paced.

Weeks one and two: analyse and say nothing. The first scan produces around eleven thousand findings, which is entirely normal for a codebase of that age and would be demoralising if presented as a work queue. It is not shared as a target. What it is used for is tuning: the ten most frequent rules account for over half the findings, and four of them conflict with conventions the team adopted deliberately years ago. Those four are disabled. Generated code and database migrations are excluded, removing another substantial share.

Week three: set the baseline and enable a minimal gate. The baseline is today. The gate has two conditions — no new bugs, no new vulnerabilities — and it blocks merging. The first week produces three failures, all genuine: a null dereference on an error path, a resource left unclosed, and a query built by string concatenation. Nobody argues with any of them, which is the point of starting narrow.

Weeks four to six: add hotspot review and duplication. Hotspot review turns out to be the most valuable addition, because it surfaces two places where certificate validation was disabled during a long-past debugging session and never restored. Neither was a finding the tool could assert was wrong; both were genuinely wrong, and a human looking at them established that in under a minute.

Weeks seven to nine: add coverage on new code. The threshold is set by looking at what the team's recent well-tested changes actually achieved, rather than at a round number. Setting it at an aspirational figure would have produced either failures nobody could meet or tests written to satisfy the gate. Set realistically, it passes most of the time and fails informatively.

Weeks ten to twelve: opportunistic cleanup. No remediation project is created. Instead, the convention becomes that when a file is being changed anyway, its existing findings are addressed. Over the quarter this clears a meaningful share of the backlog in exactly the files that matter — the ones under active development — while the untouched majority stays untouched and costs nothing.

At the end of the quarter, the total finding count has fallen by perhaps a fifth, which is not the interesting number. The interesting number is that new code has essentially no findings, the gate has never been bypassed, and nobody has asked for an exemption. That is what successful adoption looks like, and it is entirely a consequence of gating on new code rather than on the whole codebase.

17. Frequently asked questions

Does it replace code review?

No. It detects patterns known to be problematic; it cannot tell whether the code implements the right requirement, whether the design is sensible, or whether the approach fits the system. What it does is remove the mechanical findings from review entirely, so reviewers spend their attention on the things only a human can assess. That is a genuine improvement to review quality rather than a replacement for it.

What coverage threshold should we set?

Derive it from what your team's well-tested changes already achieve, rather than choosing a round number. Set it slightly above the current typical level so it is achievable but not automatic. Apply it to new code only. And watch for assertion-free tests appearing after you introduce it — that is the reliable signal that the threshold is too high or being treated as a target rather than a floor.

How do we handle a large existing backlog?

Do not treat it as a work queue. Set the baseline to today, gate on new code, and address historical findings opportunistically when touching a file for other reasons. Remediation projects targeting the backlog directly consume significant time, deliver no user-visible outcome, and are usually cancelled halfway — leaving the codebase churned and no better.

Should the same gate apply to every project?

A small number of shared gates rather than one per project, so standards are consistent and changes are made once. A stricter gate for critical services is reasonable. What fails is per-project customisation, which drifts until no two projects mean the same thing by passing, and aggregate reporting becomes meaningless.

How do we deal with false positives?

Mark them with a reason, and track the rate. Occasional false positives are inherent to the technique and acceptable. A rule producing them frequently should be disabled or have its severity reduced, because the real cost is not the individual dismissal — it is that developers begin approaching every finding expecting it to be wrong, which destroys the tool's value entirely.

Is it useful for security specifically?

For vulnerabilities in code you wrote, yes — the taint analysis genuinely finds injection risks and similar issues. It does not cover vulnerable dependencies, which are frequently the more common real attack path and need separate scanning. Nor does it find authorisation flaws, which require knowing what should be permitted. Treat it as one layer among several rather than as a security programme.

What about languages with weaker rule support?

Coverage varies considerably by language. For well-supported languages the analysis is thorough; for others it is shallower and may be better complemented by language-specific tooling feeding results in. Check the depth for your stack before assuming the same value across a polyglot estate, and be prepared to combine it with dedicated linters where its native support is thin.

What is the single decision that determines success?

Gating on new code with a baseline set to adoption day. Every failed adoption shares the same root cause — a gate applied to an entire codebase, producing an unachievable standard, resentment, exemptions and eventual abandonment. Every successful one applies strict standards to a small, fresh surface that developers can genuinely control.

Glossary

TermWhat it means
New codeCode added or changed since a defined baseline. The scope everything should be gated on.
BaselineThe date or version separating legacy code from new code. Setting it to adoption day is what makes adoption viable.
Quality gateThe conditions a project must satisfy to pass. A gate that does not block merging is a metric, not a control.
Quality profileThe set of rules active for a language. Shared centrally so a rule change is made once.
BugCode that will demonstrably behave incorrectly. The highest-value finding category.
VulnerabilityA security weakness with a plausible exploit path the tool is confident about.
Security hotspotSecurity-sensitive code requiring human judgement. Review and record a decision; do not silence.
Code smellA maintainability problem — confusing, duplicated or over-complex code.
Taint analysisFollowing untrusted input from entry to a sensitive operation. How injection risks are detected.
Cognitive complexityAn estimate of how hard code is for a human to understand, penalising nesting over sequence.
Technical debt ratioEstimated remediation effort against estimated development cost. A conversation starter, never a target.
Branch analysisAnalysing a branch against its target so a developer sees only what their change introduced.

Two of these carry the whole approach. New code and baseline together are the difference between a standard developers accept and one they resent — strict rules applied to fifty fresh lines are reasonable, and the same rules applied to fifty thousand inherited ones are not. Every failed adoption traces back to confusing the two.

Key takeaways

  • Gate on new code, not the whole codebase. This single decision determines whether adoption succeeds.
  • Tune the rule set before gating. A finding must be credible by default or it will be dismissed by default.
  • The gate must block. A reporting-only gate changes nothing.
  • Coverage is a signal, never a target. Targets produce tests that assert nothing.
  • Hotspots need review, not silencing. They are the category only a human can resolve.
  • Results belong in the pull request, where the decision is made — not in a dashboard nobody opens.

Static analysis done well is unremarkable: a small number of credible findings on the code you just wrote, fixed before merge, taking a few minutes. Done badly it is an enormous list nobody reads. The difference is almost entirely configuration, and the configuration takes an afternoon.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch