Skip to main content
Blog

Navigating the Cloud: Unleashing the Power of AWS Services

Last updated Architecture

The hardest thing about AWS is not any individual service. It is that there are hundreds of them, several will solve your problem, and the documentation for each explains what it does rather than when you should choose it over the alternatives. New teams do not struggle with complexity; they struggle with the paralysis of too many reasonable options.

This guide is a map. It covers the service categories that matter, how to choose within each, the decisions that are expensive to reverse, and the small number of practices that separate a well-run account from an alarming bill. It assumes you understand servers and databases but not this particular platform.


What you will learn
  • The mental model that makes the service catalogue navigable
  • How to choose compute, storage, database and networking without guessing
  • The account structure and identity decisions to get right before anything else
  • Where costs actually come from, and how to control them
  • Reliability patterns that use the platform properly
  • The mistakes that produce outages and surprise invoices
In this article
  1. A mental model for the catalogue
  2. Regions, zones and the failure model
  3. Accounts, identity and the foundations
  4. Networking
  5. Choosing compute
  6. Storage
  7. Databases
  8. Messaging and integration
  9. Content delivery and the edge
  10. Observability
  11. Security posture
  12. Cost: where it comes from
  13. Reliability patterns
  14. Twelve mistakes
  15. A worked example: a small production system
  16. Frequently asked questions

1. A mental model for the catalogue

The catalogue becomes navigable once you sort services by how much of the operational burden the provider absorbs. Every category — compute, storage, databases, messaging — has options along the same spectrum.

LevelYou manageThey manageTrade
InfrastructureOperating system, patching, scaling, capacityHardware, network, powerMaximum control, maximum work
ContainersContainer image, some scaling policyHost, orchestrationPortability with less operational load
Managed serviceConfiguration and dataPatching, backups, failoverLess control, far less work
ServerlessCode and configuration onlyEverything else, including capacityLeast work, most constraints

The useful default is to choose the most managed option that meets your requirements, and to move down the spectrum only when a specific constraint forces you. Teams that start at the infrastructure level because it feels safer end up operating a data centre in someone else's building, which is the worst of both arrangements.

The second organising idea: most services fall into one of eight categories — compute, storage, database, networking, messaging, identity, observability, and developer tooling. Almost everything else is a specialised variant or a higher-level assembly of these.

2. Regions, zones and the failure model

A region is a geographic area. An availability zone is one or more discrete data centres within that region, with independent power and networking, connected to other zones by low-latency links.

Two rules follow, and they drive most architecture decisions:

Spread across zones by default. Zone failures happen. Running in a single zone means a zone failure is your outage. Multi-zone deployment is usually a configuration choice with modest cost, and it is the highest-value reliability decision available.

Multi-region is a much bigger commitment. It means data replication with real consistency trade-offs, traffic routing, and doubled operational surface. Justify it with a genuine requirement — regulatory data residency, or a recovery objective that a region outage would breach — rather than with ambition. Many organisations that built multi-region architectures have never used the second region and pay for it monthly.

Region choice itself is driven by three things: where your users are, where your data may legally reside, and price, which varies meaningfully between regions. Service availability also differs — newer services appear in major regions first.

3. Accounts, identity and the foundations

This is the section to read before provisioning anything, because these decisions are expensive to change later.

Multiple accounts, not one. Separate accounts per environment and per major workload provide the strongest isolation boundary the platform offers: a mistake in development cannot touch production, and costs are attributable without tagging discipline alone. An organisation structure with centralised billing and policy makes this manageable rather than chaotic.

Federate identity. Users should sign in through your existing identity provider with short-lived credentials, not through per-account users with long-lived access keys. Static keys are the credential most likely to leak and least likely to be rotated, and they appear in most published breach post-mortems.

Roles for workloads. An application should assume a role granting exactly the permissions it needs. Embedding credentials in configuration is the pattern that keeps producing incidents, and it is entirely avoidable.

Guardrails at the organisation level. Policies that prevent whole categories of mistake — disabling logging, using unapproved regions, creating public storage — are preventive rather than detective and apply regardless of individual account permissions.

Tag from the start, enforced by policy. Owner, environment, cost centre, data classification. Without enforcement, tagging never reaches completeness, and cost allocation becomes archaeology.

4. Networking

A virtual private cloud is your isolated network. The design decisions that matter:

Address ranges planned centrally. Overlapping ranges between environments, or with an on-premises network, are painful to fix after workloads are deployed and trivial to avoid with an afternoon of planning. Leave room for growth and for future acquisitions.

Public and private subnets. Anything not requiring inbound internet access belongs in a private subnet. Load balancers sit in public subnets; application servers and databases do not. This is the single most effective network control available.

Security groups as the primary control. They are stateful firewalls attached to resources, and referencing one group from another — allowing the database group to accept connections from the application group — is far more maintainable than managing address ranges.

Private connectivity to services. Traffic to managed services can traverse the provider's network rather than the public internet, which improves both security posture and, frequently, cost.

Watch the gateway costs. Network address translation gateways charge for data processed, and a workload pulling large volumes through one accumulates charges quietly. This is a routine cause of unexplained bill increases.

5. Choosing compute

OptionBest forAvoid when
FunctionsEvent-driven work, sporadic traffic, glue between servicesLong-running jobs, sustained high volume, strict latency floors
Serverless containersStandard web services without cluster managementYou need fine-grained control over placement
Managed KubernetesMany services, existing Kubernetes expertise, portability goalsA handful of services and no existing expertise
Virtual machinesLegacy workloads, specialised licensing, unusual requirementsAnything a managed option covers
BatchLarge-scale scheduled processingInteractive or latency-sensitive work

The decision sequence that works: if the workload is event-driven and short, use functions. If it is a standard long-running service, use serverless containers. Reach for orchestration when the number of services and the sophistication of your deployment needs genuinely justify the operational weight — which is later than most teams assume. Use virtual machines when something specific requires them.

Two considerations that decide many cases. Cold starts matter for user-facing latency with functions; they are manageable with provisioned capacity but that erodes the cost advantage. And sustained load economics favour containers or instances — functions are excellent value for spiky traffic and poor value for a service that runs continuously at high volume.

6. Storage

Three fundamentally different kinds, frequently confused.

Object storage holds files addressed by key. Effectively unlimited, extremely durable, cheap, and accessible over HTTP. Use it for uploads, backups, static assets, data lake files and logs. It is not a filesystem — there are no real directories and no partial writes — and treating it as one causes performance surprises.

Block storage is a virtual disk attached to one instance. Use it for the operating system and for databases you run yourself. Performance characteristics vary by volume type, and choosing the cheapest type for a database is a common and painful mistake.

File storage provides a shared filesystem several instances can mount. Use it when applications genuinely need shared files with filesystem semantics. It is more expensive than object storage and should not be a default.

The cost lever that matters most is lifecycle policy. Object storage offers tiers from immediately available to archival, differing in price by an order of magnitude. Data older than a threshold moving automatically to cheaper tiers, and expiring entirely at the end of its retention period, is a configuration change that frequently produces the largest single storage saving available.

7. Databases

TypeUse whenWatch for
Managed relationalTransactions, joins, established schema — the default choiceVertical scaling limits; plan read replicas early
Cloud-native relationalHigher throughput and faster failover with the same interfaceHigher cost; some behavioural differences
Key-value / documentKnown access patterns, very high scale, predictable latencyAccess patterns must be designed up front; ad-hoc queries are expensive
In-memory cacheReducing read load, session storage, rate limitingIt is a cache — plan for it being empty
Data warehouseAnalytical queries over large volumesNot for transactional workloads
SearchText search, filtered faceted queriesOperationally heavier than it appears

The default should be a managed relational database unless you have a specific reason otherwise. Relational databases handle far more scale than folklore suggests, support flexible queries as requirements change, and are understood by everyone on your team. Key-value stores deliver exceptional scale and latency, in exchange for requiring you to know your access patterns in advance — which is a poor fit for a product still discovering what it needs.

Whatever you choose, three operational essentials: automated backups with a tested restore, multi-zone deployment for anything production, and encryption enabled from the start because adding it later requires a migration.

8. Messaging and integration

Decoupling components through messages is what makes distributed systems tolerable, and there are three distinct patterns.

Queues deliver each message to one consumer. Use them to smooth load, to decouple a slow operation from a fast request, and to retry failed work. Always configure a dead-letter destination for messages that repeatedly fail, or they will cycle forever and obscure the real problem.

Topics broadcast one message to many subscribers. Use them for notifications where several systems care about the same event.

Event buses route events by content to different targets, with filtering rules. Use them as the backbone when many services publish and consume events, because routing rules are configuration rather than code.

Workflow orchestration handles multi-step processes with retries, branching, error handling and long waits — expressed as a definition rather than as code holding state. This is the right tool for business processes spanning hours or days, where writing your own state machine means writing your own bugs.

The design rule that matters most: consumers must be idempotent. Delivery is at-least-once, duplicates will occur, and a consumer that charges a card or sends an email without a deduplication key will eventually do it twice.

9. Content delivery and the edge

A content delivery network caches responses close to users. For any application with a web front end, this is among the highest-return configurations available: it reduces latency, reduces load on your origin, absorbs traffic spikes, and often reduces cost because edge delivery is cheaper than origin bandwidth.

Two configuration points decide whether it works. Cache keys must include everything that changes the response and nothing that does not — including a header that varies per user means caching nothing, while omitting one that matters means serving the wrong content. And invalidation should be triggered by the events that actually change content, rather than by short expiry times that defeat the purpose.

Edge compute — small functions running at the delivery points — is useful for request manipulation, authentication checks, redirects and personalisation without a round trip to the origin. Keep this logic minimal; complex behaviour at the edge is difficult to debug and difficult to roll back.

10. Observability

The platform provides the building blocks; you must assemble them deliberately.

Logs should be structured and centralised, with a correlation identifier propagated through every component. Set retention deliberately — indefinite retention of verbose logs is a recurring and avoidable cost.

Metrics beyond the defaults. Platform metrics tell you about infrastructure; you need application metrics describing user-visible behaviour: request rate, error rate, latency distribution and queue depth.

Traces across service boundaries, which is the only practical way to answer where the time went in a request crossing several components.

Alarms on symptoms. Alert on error rates and latency for user journeys, not on individual instance health. An instance being replaced is normal; checkout failing is not.

Dashboards per service, owned by the team that runs it, showing the handful of signals that matter rather than every available metric.

11. Security posture

  • Identity first. Federated sign-in, multi-factor authentication universally, roles for workloads, no static keys. This addresses the largest share of realistic compromise.
  • Least privilege, then review. Start restrictive, widen on evidence, and schedule a review — permissions granted during an incident are permanent unless someone removes them.
  • Preventive guardrails. Policies blocking public storage, unapproved regions and disabled logging are worth more than alerts reporting the same after the fact.
  • Encryption everywhere, at rest and in transit, with managed keys unless your data classification requires you to control them.
  • Secrets in a secrets manager, injected at runtime, rotated on a schedule that has actually been exercised.
  • Scan infrastructure code before deployment. Misconfiguration is now the leading cause of cloud exposure, and it is far cheaper to catch in a pull request.
  • Understand the shared responsibility line for each service. The provider secures the platform; you secure configuration, identity, data and access — and the line moves depending on how managed the service is.

12. Cost: where it comes from

SourceControl
Over-provisioned computeRightsize on measured utilisation after weeks of real traffic
Non-production running constantlySchedule development and test environments off outside working hours
On-demand pricing for steady loadCommit to discounted terms — but only after rightsizing
Data transferKeep chatty components in the same zone; watch gateway processing charges
Storage growthLifecycle policies to cheaper tiers, and expiry at end of retention
Orphaned resourcesEnforced tagging plus automated cleanup of unattached volumes and idle addresses
Log retentionSet it deliberately; verbose logs kept forever are a large silent cost
Idle managed servicesProvisioned clusters nobody uses continue billing

Two practices matter more than any individual optimisation. Make cost visible per team, broken down by service and tagged owner, reviewed monthly — teams optimise what they can see, and central cost exercises without visibility produce savings that erode within two quarters. And set budget alerts and hard limits, because a misconfigured job or a runaway query will happen, and the difference between noticing on day one and on day thirty is substantial.

13. Reliability patterns

Multi-zone by default for anything production. This is the highest-value reliability decision and usually the cheapest.

Health checks that check the right thing. A check confirming the process is running will keep routing traffic to an instance that cannot reach its database. Check the dependencies that determine whether the instance can actually serve.

Automatic scaling on a metric that reflects load — request count or queue depth rather than processor utilisation, which frequently lags the thing you care about.

Timeouts, retries with backoff and jitter, and circuit breakers on every call between components. Synchronised immediate retries turn a brief degradation into an outage.

Graceful degradation. If a recommendation service is unavailable, the page should render without recommendations rather than failing. Decide in advance which dependencies are genuinely mandatory.

Tested restores. Backups that have never been restored are an assumption. Restore into a clean environment quarterly and record how long it took — that number is your real recovery objective regardless of what any document claims.

14. Twelve mistakes

  1. One account for everything. No isolation, no cost attribution, no blast radius control.
  2. Long-lived access keys. The credential most likely to leak.
  3. Single-zone deployment. A zone failure becomes your outage.
  4. Databases in public subnets. Reachable from the internet by configuration error.
  5. No tagging discipline. Costs unattributable, resources unowned, cleanup impossible.
  6. Committing to discounted pricing before rightsizing. Locking in the wrong capacity for a year.
  7. Kubernetes for three services. Substantial operational weight for no benefit.
  8. Choosing a key-value store before knowing access patterns. Expensive to change afterwards.
  9. Non-idempotent message consumers. Duplicate charges and duplicate emails.
  10. No dead-letter destination. Poison messages cycling forever, obscuring the real failure.
  11. Health checks that only check the process. Traffic routed to instances that cannot serve.
  12. Untested backups. Discovered during the incident they were meant to solve.

15. A worked example: a small production system

Consider a straightforward requirement: a web application with an API, a database, background jobs, file uploads and a modest amount of traffic. This is the most common shape, and it is worth walking through because the choices are representative.

Accounts and identity come first. Three accounts — development, staging, production — under one organisation with centralised billing and a small set of preventive policies. Engineers sign in through the existing identity provider and assume roles; nobody has a static key. This takes a couple of days and is the foundation everything else assumes.

Networking is deliberately simple. One virtual network per account, spanning three zones, with public subnets containing only the load balancer and private subnets containing everything else. Address ranges are planned across all three accounts at once, which costs an hour and prevents a class of problem that is miserable to fix later.

Compute is serverless containers. The application is a standard long-running web service, so functions are a poor fit and Kubernetes is unwarranted for a handful of services. Containers behind a load balancer, scaling on request count, deployed from a pipeline that builds the image once and promotes the same image through environments.

The database is managed relational, multi-zone. Not because relational is fashionable but because the access patterns are still evolving and flexible querying is worth more at this stage than the scale a key-value store would offer. Automated backups, a read replica added when reporting queries start affecting the application, and encryption enabled from creation.

Uploads go to object storage, never to the container filesystem, with a lifecycle policy moving files to a cheaper tier after ninety days and expiring them at the end of the retention period. Uploads are signed directly from the browser so large files never traverse the application.

Background work goes through a queue. The web request enqueues and returns; a separate consumer processes. Consumers are idempotent, keyed on a message identifier, with a dead-letter destination configured on day one rather than after the first poison message.

Delivery sits behind a content delivery network, with static assets cached aggressively and the cache key carefully scoped so that per-user headers do not defeat caching entirely.

Observability is assembled deliberately. Structured logs with a correlation identifier, application metrics for request rate, errors and latency, traces across the queue boundary, and alarms on user-visible symptoms rather than instance health. Log retention is set to ninety days rather than left indefinite.

Cost controls exist before the first invoice. Enforced tagging, a monthly per-service review, budget alerts, and development environments scheduled off outside working hours. Rightsizing happens after four weeks of real traffic, and only then is any discounted commitment considered.

None of this is exotic, and that is the point. The great majority of well-run systems on this platform look approximately like the above, and the teams that struggle are usually those that reached for orchestration, multi-region or a specialised database before the requirements justified it.

16. Frequently asked questions

Where should a team new to the platform start?

Account structure, identity federation and infrastructure as code — before deploying anything. These are the decisions most expensive to retrofit, and a single-account, console-clicked environment becomes progressively harder to correct. Two weeks on foundations saves considerably more later.

Serverless or containers?

Functions for event-driven and sporadic work; containers for standard long-running services. The economics reverse at sustained volume, where functions become expensive relative to a continuously running container. Many systems use both, and that is a reasonable outcome rather than an inconsistency — choose per workload rather than adopting one philosophy for everything.

Do we need Kubernetes?

You need deployment, service discovery, health checking and scaling. Kubernetes provides all of that comprehensively at the cost of real operational complexity. Below roughly ten to fifteen services, serverless container platforms deliver the same outcome with far less to learn and maintain. Adopt it when the number of services or the sophistication of your deployment needs genuinely justify it, or when portability across providers is a firm requirement.

How do we avoid vendor lock-in?

Partly by accepting some of it. Using managed services is where most of the value is, and refusing them to preserve portability means operating everything yourself — which is a certain cost paid against an uncertain benefit. Practical mitigation: keep business logic free of provider-specific concerns, use open formats for stored data, define infrastructure as code, and know what migrating each component would actually involve. That is usually enough leverage without sacrificing the platform's advantages.

What causes surprise bills?

In rough order: data transfer through gateways, forgotten resources in unused regions, log retention set to indefinite, non-production environments running continuously, and runaway processes such as a recursive trigger or an unbounded query. Budget alerts catch the acute cases; monthly per-service review catches the gradual ones, which are usually larger in total.

How many environments do we need?

At minimum development, staging and production in separate accounts. Ephemeral per-branch environments are increasingly practical with infrastructure as code and remove the shared staging bottleneck that slows most teams. What matters is that non-production environments are defined the same way as production, differing only in parameters — environments that have drifted apart make staging results meaningless.

Should we use a multi-cloud strategy?

Rarely by choice at the outset. It multiplies operational surface, skills required and security configuration, in exchange for negotiating leverage most organisations never exercise. Legitimate drivers exist — regulation, an acquisition, a capability available only elsewhere — but adopting it pre-emptively usually means running two platforms badly instead of one well.

What is the most valuable practice to adopt early?

Infrastructure as code, applied to everything including identity and networking. It makes environments reproducible, changes reviewable, and recovery possible from definitions rather than memory. It also makes every other practice easier to enforce — guardrails, tagging and consistency all become properties of the code rather than of individual discipline.

A foundations checklist

Before a workload goes to production, these should all be true. None of them are difficult; all of them are expensive to retrofit across a deployed estate.

ItemWhy it matters
Separate accounts per environmentThe strongest isolation boundary available, and the basis of cost attribution
Federated sign-in, no static keysRemoves the credential most likely to leak
Roles for workloadsNo credentials embedded in configuration
Address ranges planned centrallyOverlaps are painful once workloads exist
Databases in private subnetsNot reachable from the internet by misconfiguration
Multi-zone for anything productionA zone failure stops being your outage
Infrastructure defined as codeReviewable, reproducible, recoverable
Tagging enforced by policyRetrospective tagging never completes
Preventive guardrails in placeBlocking a mistake beats reporting it
Budget alerts configuredRunaway costs noticed on day one, not day thirty
Backups restored and timedAn untested backup is an assumption
Alarms on user-visible symptomsInstance health is not what customers experience

Teams that complete this list before their first production deployment rarely have foundational problems later. Teams that defer it spend a subsequent quarter retrofitting isolation and identity across an estate that was never designed for either.

Key takeaways

  • Choose the most managed option that meets your requirements. Move down the spectrum only when forced.
  • Get accounts, identity and networking right first. These are the expensive things to retrofit.
  • Multi-zone by default; multi-region only with a real requirement.
  • Managed relational is the right database default until a specific constraint says otherwise.
  • Idempotent consumers and dead-letter destinations are not optional in message-driven systems.
  • Make cost visible per team. Rightsize before committing, and set budget alerts before the first invoice.

The teams that do well on this platform are not the ones using the most services. They are the ones who chose a small number of well-understood building blocks, defined them as code, made the costs visible, and left themselves room to change their minds later.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch