Skip to main content
Blog

Cloud Migration: A Practical Strategy Guide

Last updated AWS

Cloud migration has a reputation problem. It is sold as a technology project and experienced as an operating-model change, which is why so many programmes deliver the infrastructure and none of the benefits. The servers move, the bill arrives larger than before, the teams work the same way they always did, and somebody asks what exactly was gained.

This guide treats migration as what it actually is: a sequence of decisions about which workloads move, in what form, in what order, and what changes about how you run them afterwards. It covers assessment, the realistic migration paths, cutover mechanics, data movement, cost control, and the organisational work that determines whether the investment pays back.


What you will learn
  • How to decide whether migration is justified, and what the honest benefits are
  • A portfolio assessment that produces decisions rather than a spreadsheet
  • The six migration paths and how to choose between them per workload
  • Data migration and cutover — the parts that actually go wrong
  • Landing zones, networking, identity and security foundations
  • Why costs rise after migration, and the practices that bring them down
In this article
  1. Why migrate, honestly
  2. Assessment that produces decisions
  3. The six migration paths
  4. Choosing a path per workload
  5. The landing zone
  6. Sequencing the programme
  7. Data migration
  8. Cutover mechanics
  9. Networking and hybrid connectivity
  10. Security and identity
  11. Cost: why it goes up, and how to bring it down
  12. Operating model after migration
  13. Twelve things that go wrong
  14. A worked example: migrating one dependency cluster
  15. Frequently asked questions

1. Why migrate, honestly

The benefits are real, but they are not automatic, and being specific about which ones you are pursuing shapes every subsequent decision.

Claimed benefitReality
Lower costOnly with rightsizing, elasticity and discipline. A lift-and-shift of steady workloads usually costs more.
ElasticityGenuine and valuable — but only for workloads whose demand actually varies
Speed of deliveryThe largest real benefit: environments in minutes rather than months
ReliabilityBetter building blocks; your architecture still determines the outcome
SecurityBetter primitives and worse defaults — misconfiguration is now the main risk
Access to managed servicesFrequently the strongest argument: databases, queues and analytics you no longer operate
Exiting a data centreA hard deadline that concentrates minds, and often the true driver

Be explicit about which of these justifies the programme, because they conflict. Optimising for speed of exit means lifting and shifting. Optimising for cost means re-architecting. Attempting both at once produces a programme that does neither well and runs a year long.

It is also legitimate to conclude that some workloads should not move. A stable, heavily utilised system on hardware that is paid for, with no elasticity requirement and a team that knows it well, may be more expensive to run in the cloud with no offsetting benefit. Saying so early is a sign of a well-run programme, not a failure of ambition.

2. Assessment that produces decisions

Most assessments produce an inventory. An inventory is not a decision. What you need for each workload is enough information to choose a path, and that requires six facts.

  1. Business criticality and acceptable downtime. Determines cutover strategy and how much rehearsal is justified.
  2. Dependencies, in both directions. What it calls, and what calls it. This is the single most commonly underestimated item, and the usual source of migration-day surprises.
  3. Data volume, sensitivity and residency. Determines transfer method and whether some options are legally excluded.
  4. Utilisation profile. Steady or spiky, and what the actual usage is rather than the provisioned capacity. Most on-premises servers are dramatically over-provisioned.
  5. Change rate and ownership. Actively developed workloads can be re-architected; frozen ones with no owner cannot.
  6. Technical constraints. Operating system versions, licences, hardware dependencies, hard-coded addresses.

Automated discovery tooling helps with inventory and network flows, and is worth using. It will not tell you which system nobody dares touch, or which nightly job the finance team depends on and nobody documented. Combine tooling with conversations, and treat any workload whose owner cannot be identified as a risk rather than as a simple case.

The dependency map is the deliverable

If the assessment produces one artefact worth the effort, it is an accurate map of what talks to what. Migration order, cutover grouping, network design and rollback planning all derive from it. Programmes that skip it discover their dependencies during cutover weekends, which is the most expensive possible moment.

3. The six migration paths

PathWhat it meansEffortBenefit realised
RetireTurn it off; nobody uses itMinimalImmediate cost and risk reduction
RetainLeave it where it is, for nowNoneNone — but avoids wasted effort
RehostMove the machine as-isLowExit the data centre; little else
ReplatformMove with targeted changes — managed database, managed load balancerMediumReduced operational burden, some cost benefit
RepurchaseReplace with a SaaS productMedium, mostly non-technicalEliminates the workload entirely
RefactorRe-architect for cloud-native patternsHighFull elasticity, resilience and cost benefit

Two observations that experienced programmes converge on. First, retire is underused. A serious inventory typically finds between ten and thirty percent of workloads that nobody needs — decommissioned projects, duplicated reporting, environments provisioned for a project that ended. Every one you retire is a workload you do not migrate, test, secure or pay for.

Second, replatform is the sweet spot for most workloads. Moving a self-managed database to a managed one, or a hand-built load balancer to a managed service, captures a large share of the operational benefit for a fraction of the effort of refactoring. Full refactoring should be reserved for workloads where the business case is specific and strong.

4. Choosing a path per workload

A short decision sequence resolves most cases without lengthy debate.

Is anyone using it?

Check actual traffic and logins, not opinion. If not: retire, after a notice period and a snapshot.

Does a mature product do this?

Email, file sharing, ticketing, HR, CRM, monitoring. Repurchasing removes the workload permanently rather than relocating it. The obstacles are usually data migration and process change rather than technology.

Is it actively developed by a team that owns it?

If yes, replatform or refactor are viable — someone can absorb the change. If it is frozen with no owner, rehost or retain; re-architecting code nobody understands is how migration programmes lose a year.

Does demand vary significantly?

Elasticity is where cloud economics genuinely work. A workload that is busy for four hours a day or has strong seasonal peaks justifies real effort. One at steady seventy percent utilisation will simply cost more without redesign.

What is the deadline?

A hard data-centre exit date makes rehosting the pragmatic answer for the long tail, with optimisation planned deliberately afterwards. That is a legitimate strategy provided the second phase is actually funded — the failure mode is a programme that declares victory at rehost and never returns.

5. The landing zone

The landing zone is the foundation everything lands on, and it must exist before the first workload moves. Retrofitting it across two hundred deployed workloads is an order of magnitude harder than building it first.

Six elements, all of which should be defined as code from the outset:

  • Account and subscription structure. Separation by environment and by business domain, so a mistake in development cannot affect production and costs are attributable without guesswork.
  • Identity. Single sign-on from your existing directory, role-based access, no long-lived credentials, and break-glass accounts that are documented and monitored.
  • Network topology. Address ranges planned to avoid collisions with on-premises and with future acquisitions, connectivity to existing sites, and a clear model for what may reach what.
  • Guardrails. Preventive policies that make the wrong thing impossible — public storage buckets blocked, unapproved regions denied, encryption required — rather than detective controls that report violations afterwards.
  • Logging and monitoring. Centralised, tamper-resistant, retained appropriately, and configured before there is anything to log.
  • Tagging standard. Owner, environment, cost centre, data classification. Enforced at creation, because retrospective tagging never completes.

The tagging standard sounds trivial and is disproportionately consequential. Without it, cost allocation is impossible, ownership is unknown, and the cleanup of orphaned resources becomes archaeology. Enforce it with policy so untagged resources simply cannot be created.

6. Sequencing the programme

Order the programme to build capability and confidence before touching anything that matters.

WaveWhat movesPurpose
0Landing zone, connectivity, identity, pipelinesFoundations, before any workload
1Two or three low-risk, low-dependency workloadsProve the process; find the unknown unknowns
2Internal applications with contained blast radiusBuild team capability at real scale
3Customer-facing systems, grouped by dependency clusterThe main body of work
4The hard tail: legacy, licensed, tightly coupledSlowest per workload; needs the most time
5Optimisation and decommissioningWhere cost benefits are actually realised

Two sequencing rules save more time than any tool. Migrate dependency clusters together — systems that chat constantly should not be split across a network boundary for months, because the latency and failure modes will consume the schedule. And never start with the hardest system, however tempting it is to prove the concept: an early failure on a critical workload can end the programme's political support entirely.

Wave 5 is the one that gets cut when the programme runs late, and it is the wave where most of the financial benefit lives. Fund and schedule it explicitly from the beginning.

7. Data migration

Data is where migrations overrun. Three questions determine the approach.

How much, and how fast can it move?

Compute the transfer time honestly against your actual available bandwidth, not the theoretical link speed, and include the fact that the link is also carrying production traffic. Beyond a certain volume, physical transfer appliances are faster than the network — a conclusion that surprises people until they do the arithmetic.

How much downtime is acceptable?

This drives everything. A workload tolerating a weekend can be migrated by backup and restore, which is simple and reliable. A workload requiring minutes needs continuous replication with a short final cutover. A workload requiring zero downtime needs dual-write or read-replica promotion, which is significantly more complex and needs proportionate rehearsal.

How will you prove it arrived correctly?

The step most often skipped. Row counts by table, checksums on critical columns, aggregate totals compared against the source, and a sample of records compared field by field. Automate the comparison and run it after every rehearsal. "It looked fine" is not a validation strategy, and discovering a truncated column three weeks later is considerably worse than discovering it during a rehearsal.

Practical guidance

  • Migrate a copy first and run the new environment in parallel with production traffic mirrored, if you can.
  • Freeze schema changes during the migration window; a schema change mid-replication is a reliable way to break it.
  • Plan for the source remaining authoritative until validation completes, and keep it available afterwards for longer than feels necessary.
  • Test the rollback path with real data volumes, not a sample.

8. Cutover mechanics

Cutover is a rehearsed procedure, not an event to be improvised. Everything below should be written down and practised before the real attempt.

  1. Entry criteria. What must be true before starting: validation passed, monitoring live, team available, change approved, communications sent.
  2. The runbook. Every step, with an owner and expected duration. Long enough to be unambiguous, short enough that someone can follow it at three in the morning.
  3. Go/no-go checkpoints. Defined points where the team explicitly decides to continue or roll back, with named decision-makers.
  4. The rollback plan. Written before starting, with a point of no return clearly identified — usually the moment the new system accepts writes that the old one will not have.
  5. Verification. Specific checks proving the system works, not merely that it starts: a real transaction end to end, key reports producing correct figures, integrations exchanging messages.
  6. Hypercare. A defined period of elevated monitoring and staffing afterwards, with clear criteria for declaring the migration complete.

Two techniques significantly reduce risk. Rehearse in full against a production-scale copy, timing each step — the first rehearsal always reveals a missing step, and the second usually reveals another. And reduce DNS time-to-live values well in advance, so that when you switch, traffic actually follows within minutes rather than hours.

9. Networking and hybrid connectivity

Networking is where migration programmes discover that the diagram was wrong. Plan for a hybrid period lasting longer than intended, because it always does.

Address space. Plan ranges centrally with room to grow. Overlapping ranges between cloud and on-premises, or between environments, are painful to fix once workloads are deployed and are entirely avoidable with an afternoon of planning.

Connectivity. A dedicated private link offers predictable latency and throughput but takes weeks to provision — order it early. A site-to-site VPN is fast to establish and adequate for many workloads; run one as a backup path regardless.

Latency. Applications that made a hundred small calls to a database in the same rack behave very differently when that database is fifteen milliseconds away. This is the most common cause of "it worked in testing and crawls in production", and it is why dependency clusters should move together.

Data transfer costs. Traffic leaving the cloud is charged, and it is easy to build an architecture that shuttles data back and forth across the boundary. Model this during design rather than discovering it on the first full month's invoice.

10. Security and identity

The cloud provides better security primitives and worse defaults than a private data centre where the network perimeter did much of the work implicitly. The dominant risk shifts from intrusion to misconfiguration.

  • Identity is the new perimeter. Federate to your existing directory, require multi-factor authentication universally, use short-lived credentials for workloads, and eliminate static access keys. Most publicised cloud breaches trace back to a credential that should not have existed.
  • Least privilege, then review. Start restrictive and widen on evidence. Permissions granted during a migration crisis are permanent unless someone removes them, so schedule a review.
  • Preventive over detective. A policy that blocks public storage is worth more than an alert reporting that public storage was created yesterday.
  • Encrypt by default, at rest and in transit, with keys you control where the data classification requires it.
  • Scan infrastructure code before it deploys. Configuration errors are far cheaper to catch in a pull request than in production.
  • Understand the shared responsibility boundary for each service you use. The provider secures the platform; you secure configuration, identity, data and access. The line moves depending on how managed the service is.

11. Cost: why it goes up, and how to bring it down

Bills rising after migration is the norm, not the exception, and the causes are consistent.

CauseFix
Provisioned to match old hardwareRightsize on measured utilisation after a few weeks of real traffic
Everything running all the timeSchedule non-production environments off outside working hours
On-demand pricing for steady workloadsCommit to discounted terms once usage is understood
Orphaned resourcesEnforced tagging plus automated cleanup of untagged and unattached resources
Unmanaged storage growthLifecycle policies moving old data to cheaper tiers, and deleting what is past retention
Cross-region and egress trafficArchitectural review; keep chatty components together
Nobody sees the costPer-team cost visibility and budget alerts

The most effective single practice is making cost visible to the teams who create it, broken down by service and tagged owner, reviewed monthly. Teams optimise what they can see. Central cost-reduction exercises without that visibility produce a one-off saving that erodes within two quarters.

Sequence matters: rightsize before committing to discounted pricing, or you will lock in the wrong capacity for a year. Wait for a few weeks of genuine production traffic before making either decision.

12. Operating model after migration

This is the part that determines whether the programme delivers anything beyond a change of address, and it is the part most often left unfunded.

Infrastructure becomes code. Manual console changes are the enemy of reproducibility. Everything defined in version control, reviewed and applied through a pipeline. Environments become disposable, which is where much of the delivery-speed benefit actually comes from.

Teams own their infrastructure. The traditional split — developers write code, an operations team runs it — reintroduces the queue that migration was supposed to remove. Give product teams the ability to provision within guardrails, with a platform team providing paved paths rather than acting as a ticket desk.

Reliability is engineered explicitly. Managed services are more reliable than most self-run equivalents, but availability still depends on your architecture: multi-zone deployment, health checks, graceful degradation, tested failover. Define service level objectives and alert on those rather than on infrastructure metrics.

Skills need real investment. The most common post-migration problem is a team operating an unfamiliar platform under pressure. Budget training time before cutover, not after, and accept that the first months will be slower while people learn.

13. Twelve things that go wrong

  1. Migrating before the landing zone is ready. Every subsequent workload inherits the gaps.
  2. Incomplete dependency mapping. Discovered at cutover, at the worst possible time.
  3. Lift and shift with no optimisation phase. Costs rise, benefits do not arrive, credibility is spent.
  4. Underestimating data transfer time. A weekend window that needed nine days.
  5. No validation of migrated data. Silent corruption found weeks later.
  6. Chatty applications split across the network boundary. Latency destroys performance.
  7. No tagging discipline. Costs unattributable, resources unowned.
  8. Permissions granted in a crisis and never revoked. A permanent security debt.
  9. Committing to discounted pricing before rightsizing. Locking in the wrong capacity.
  10. Keeping the old operating model. The same tickets and queues, on someone else's servers.
  11. Not decommissioning the source. Paying for both indefinitely.
  12. Treating it as an infrastructure project. No application team engagement, no behaviour change, no benefit.

14. A worked example: migrating one dependency cluster

Abstract waves are easy to nod along to, so consider one realistic cluster: an order management application, its database, a reporting service that reads from a replica, a nightly batch job that produces settlement files, and an integration with a third-party carrier. Five components, one business process, and a migration that has to happen as a unit because the pieces talk to each other constantly.

Assessment finds three surprises, which is typical. The reporting service turns out to be used by finance for a month-end process nobody had documented, which raises its criticality from low to high. The batch job writes files to a network share that two other systems read, an interaction that appeared nowhere in the architecture diagram. And the carrier integration authenticates by source IP address, which means the migration requires coordination with an external party who works to their own timescales — discovered in week two rather than on cutover night, which is the entire value of doing the assessment properly.

The path chosen is replatform rather than rehost. The application moves largely unchanged onto managed compute, but the self-managed database becomes a managed one, and the network share becomes managed object storage. Those two substitutions remove the two components the operations team spent most of its time on. Full refactoring is rejected: the team that owns the application is small, the workload is steady rather than spiky, and the business case for elasticity is weak.

Data migration is rehearsed three times. The first rehearsal reveals that the transfer takes eleven hours rather than the estimated four, because the available bandwidth is shared with production traffic during the day. The second reveals a character-encoding difference that silently mangles a small number of customer names — caught only because the validation compares a sample of records field by field rather than merely counting rows. The third rehearsal is clean, and it also establishes the real duration of each step, which is what makes the cutover runbook credible.

Cutover is a six-hour window on a Saturday. Continuous replication runs for the preceding week, so the final synchronisation takes minutes rather than hours. DNS time-to-live values were reduced a week earlier. There is a defined point of no return — the moment the new database accepts its first write — and a rollback plan that is genuinely usable before it. Verification is not "the application starts" but a real order placed end to end, a settlement file generated and compared against the previous day's, and a test message exchanged with the carrier.

Hypercare runs for two weeks, after which the source environment is powered off but retained for a further month before deletion. That last step is the one most commonly forgotten, and forgetting it means paying for both estates indefinitely while gradually losing the confidence to switch either off.

15. Frequently asked questions

How long should a migration take?

For a mid-sized estate of a few hundred workloads, twelve to twenty-four months is typical, with foundations taking the first two to three. Programmes promising six months usually mean rehosting a subset and deferring everything difficult. The realistic constraint is rarely technical capacity; it is the availability of the application teams who must test and accept each workload.

Should we go multi-cloud?

Rarely by choice at the start. Multi-cloud multiplies the operational surface, the skills required and the security configuration, in exchange for negotiating leverage most organisations never exercise. It makes sense when regulation demands it, when an acquisition brings a second provider, or when a specific capability exists only elsewhere. Portability is better served by keeping your architecture reasonably standard than by refusing to use managed services.

Is refactoring worth it, or should we just rehost?

Per workload rather than as a policy. Refactor where demand varies significantly, where a team actively owns the code, and where the operational burden is currently high. Rehost the long tail, especially under a deadline. The costly mistake is refactoring everything, which extends the programme by a year for benefits that only a minority of workloads will realise.

How do we handle software licensing?

Investigate early, because licence terms can make an otherwise obvious plan uneconomic. Some vendors restrict deployment on shared infrastructure, some charge differently by processor count, and some offer bring-your-own-licence arrangements that materially change the maths. This is one of the few areas where the answer genuinely requires legal and procurement involvement rather than engineering judgement.

What about workloads that legally cannot leave the country?

Most major providers operate regions in many jurisdictions, which resolves residency for a large share of cases — but confirm where data is processed as well as stored, including backups, logs and support access. Where regulation is stricter still, a hybrid model keeping the regulated data on-premises while migrating everything else is a legitimate and common outcome.

How do we keep the business running during a long migration?

Accept the hybrid period as a designed state rather than an inconvenience: invest in reliable connectivity, consistent identity across both estates, and monitoring that covers everything in one place. Keep dependency clusters intact so the boundary falls in quiet places, and set the expectation that the hybrid period will last longer than planned — because it will.

Who should run the programme?

A small central team owning the landing zone, guardrails, tooling and standards, with migration executed by the teams who own each application. Fully centralised programmes stall because the central team cannot test or accept workloads on behalf of the business. Fully devolved programmes produce inconsistent foundations that become a security and cost problem within a year.

What is the first thing to do?

Two things in parallel: build the landing zone, and produce an honest dependency map. Both take weeks, both are prerequisites for every decision that follows, and both are routinely skipped in favour of migrating something quickly to demonstrate progress. That early demonstration usually costs more time later than it saves.

Key takeaways

  • Name the benefit you are pursuing. Speed of exit and cost optimisation demand different strategies.
  • The dependency map is the deliverable. Migration order, cutover grouping and network design all derive from it.
  • Retire more than you expect. Ten to thirty percent of most estates should never be migrated at all.
  • Replatform is usually the sweet spot. Most of the operational benefit for a fraction of the refactoring effort.
  • Landing zone first. Retrofitting foundations across a migrated estate is an order of magnitude harder.
  • Costs rise unless you fund the optimisation phase. Rightsize, schedule, commit, tag, and make cost visible per team.

A migration succeeds when the teams that own the workloads can provision, deploy and recover them faster than they could before. If that is not true afterwards, the programme moved the servers and left the constraints behind.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch