Most teams use a small fraction of what their cloud provider offers, and the fraction they use is the obvious part — virtual machines, object storage, a managed database, a load balancer. That is a perfectly good cloud deployment, and it also means a fair amount of what those teams build by hand already exists as a service they are paying nothing extra to ignore.
This article covers the services past the obvious ones: the ones that quietly solve problems teams keep rebuilding, the ones with real trade-offs worth understanding before adopting, and the ones whose value only becomes visible at a particular scale. It is organised by problem rather than by product name, because that is how the decision actually gets made.
What you will learn
- Which common in-house builds have a managed equivalent
- Event-driven and workflow services, and when they fit
- Data services beyond a relational database
- Networking and security capabilities most teams under-use
- Cost control that actually works
- How to evaluate a managed service against building it
- The build-versus-adopt question
- Beyond virtual machines: the compute spectrum
- Events and messaging
- Orchestrating multi-step work
- Beyond a relational database
- Querying data where it sits
- Streaming data
- Search
- Caching
- Identity for your users
- Secrets and configuration
- Networking you probably under-use
- Content delivery and the edge
- Observability
- Governance and guardrails
- Cost control that works
- Migration and modernisation services
- How to decide
- Twelve mistakes
- A worked example: replacing three in-house systems
- Frequently asked questions
1. The build-versus-adopt question
Every managed service is the same trade: less code you own, more dependency on a provider, and a cost structure that may or may not suit you.
The arguments for adopting are strong and usually underweighted. You do not maintain it, patch it, scale it or get paged for it. It integrates with the identity and monitoring you already use. And the operational knowledge required is documentation rather than tribal.
The arguments against are real and usually overstated. Lock-in exists and is proportional to how deeply the service's model shapes your code. Cost can exceed a self-hosted equivalent at high, steady volume. And you inherit the service's limits, which occasionally do not fit.
The framing worth using: what is the operational cost of running this ourselves, honestly, including the person who gets paged? Teams consistently underestimate this, because the cost of self-hosting is spread across many people's time and does not appear on an invoice. A managed service that looks expensive next to a compute instance is frequently cheaper than the same thing plus the fraction of an engineer it consumes.
2. Beyond virtual machines: the compute spectrum
Most workloads default to virtual machines or containers, and there is a spectrum worth knowing.
| Model | You manage | Suits | Watch for |
|---|---|---|---|
| Virtual machines | OS, runtime, scaling, patching | Legacy workloads, unusual requirements | Most operational burden |
| Managed containers | Images and configuration | Most services | Cluster complexity if self-managing the orchestrator |
| Serverless containers | Images only | Variable load; teams without cluster expertise | Cold starts; cost at steady high load |
| Functions | Code only | Event handlers, glue, sporadic work | Duration limits; cold starts; local testing |
| Batch | Job definitions | Scheduled heavy compute | Overkill for small jobs |
The two under-used options are serverless containers and batch.
Serverless containers remove cluster management entirely while keeping the container model — you supply an image and it runs, scaling to demand. For teams running a handful of services without a platform team, this eliminates a large amount of work that produces no product value.
Batch handles the case of "run this heavy computation on a schedule or on demand, with the right amount of hardware, and go away afterwards" — which teams routinely implement as a permanently-running instance that is idle most of the time. Spot capacity makes this dramatically cheaper for anything interruptible.
On spot capacity generally: heavily discounted compute that can be reclaimed with brief notice. For anything fault-tolerant — batch processing, CI, stateless workers, non-urgent training — the saving is large and the engineering required is modest. Teams that have never tried it usually assume it is harder than it is.
3. Events and messaging
Three distinct patterns, frequently conflated, and choosing wrong produces awkward systems.
Queues for work distribution. A producer puts a message on, one consumer takes it off and processes it. Decouples producers from consumers, absorbs bursts, and provides retries and a dead-letter destination for failures. If your application has a background job table in the database, this is what it should be instead.
Topics for fan-out. One message, many independent subscribers. When several things must happen in response to one event, this removes the need for the producer to know about any of them.
Event buses for routing. Events from many sources, routed by rules to many targets, with filtering and transformation. This is the one most teams do not know exists and the one that most often replaces a pile of bespoke integration code — including events emitted by the cloud platform itself, which lets you react to infrastructure changes without polling.
The pattern worth internalising: a dead-letter destination is not optional. Messages that fail repeatedly must go somewhere inspectable, and a queue without one silently discards work. Every serious deployment eventually discovers this the hard way.
4. Orchestrating multi-step work
A genuinely under-used category. When a process has several steps with branching, retries, waits and error handling, the usual approach is code that coordinates it — and that code accumulates state management, retry logic and failure handling until it is the most fragile part of the system.
A workflow service handles that coordination: you define the steps and transitions, and the service manages state, retries with backoff, timeouts, parallel branches, error paths and human approval steps. Each execution is visible, with a record of what happened at every step.
Where this earns its place: order processing, onboarding flows, data pipelines with dependencies, anything with a human approval in the middle, and any long-running process where a machine cannot simply hold the state in memory for three days.
The property that sells it once you have used it is visibility. When a workflow fails at step seven of eleven, you can see exactly which execution, which step, what the input was and what the error said — and resume from there. Debugging equivalent bespoke code means reading logs and inferring.
5. Beyond a relational database
A managed relational database is the right default for most applications. Several alternatives are worth knowing for the cases where it is not.
Wide-column and key-value stores for very high throughput with simple access patterns, where the query is always by key and horizontal scale matters more than query flexibility. The cost is that you must know your access patterns in advance, because the data model is designed around them rather than around the entities.
Serverless relational databases that scale capacity with load and can drop to near zero when idle. Excellent for development environments, intermittent workloads and anything with sharply varying traffic. Frequently the right answer for the six non-production databases running at full cost.
In-memory data stores beyond caching — leaderboards, rate limiting, session state, transient computation. Fast enough to change what is architecturally feasible.
Time-series stores for metrics, telemetry and anything measured repeatedly. Purpose-built for the write-heavy, time-bounded-query pattern that general databases handle poorly, and the difference at volume is substantial.
Graph stores for relationship-heavy queries — recommendations, fraud rings, network analysis — where the equivalent relational query is a chain of self-joins that becomes unusable at depth.
Ledger and append-only stores for anything requiring a verifiable, immutable history, where the audit trail is the product rather than a byproduct.
6. Querying data where it sits
One of the highest-value under-used capabilities: query object storage directly, without loading anything into a database.
Point a query service at files in object storage, define a schema, and query them. Serverless, priced by data scanned, no infrastructure.
Why this matters: an enormous amount of data is already sitting in object storage — logs, exports, archives, event data — and the assumption is that querying it requires a pipeline into a warehouse. For exploratory analysis, occasional reporting and investigation, querying in place is faster to set up and cheaper to run.
Two practices make the difference between cheap and expensive. Use a columnar format rather than raw text, which reduces data scanned by an order of magnitude on typical queries. And partition by the columns you filter on, usually date, so a query for one day does not scan a year.
A related capability worth knowing: filtering at the storage layer, so a request for a subset of a large object returns only the matching rows rather than transferring the whole file for the application to filter. Substantially reduces both transfer cost and application memory.
7. Streaming data
When data arrives continuously and must be processed as it arrives, streaming services provide durable, ordered, replayable streams that multiple consumers read independently at their own pace.
The property that distinguishes a stream from a queue is replay. A queue message is consumed and gone. A stream retains data for a period, so a new consumer can process history, and a consumer that had a bug can reprocess after being fixed. That single difference is why streams suit analytics and event sourcing where queues suit work distribution.
Managed options cover ingestion, delivery to storage with buffering and format conversion, and processing with windowed aggregation. The delivery service in particular replaces a common in-house build: reliably getting a high volume of events into object storage in a queryable format, batched sensibly and partitioned by time.
8. Search
Full-text search implemented against a relational database with pattern matching is a well-trodden path to disappointment. It does not rank sensibly, does not handle misspellings, does not do synonyms and does not scale.
Managed search services provide relevance ranking, faceting, typo tolerance, synonyms and language handling. The choice is roughly between a managed version of a general search engine — powerful, flexible, requires understanding it — and a simpler search service that handles more for you at the cost of control.
The guidance: if search is a core part of the product experience, use a real search service. If it is a filter box on an admin page, your database is fine. The mistake is the middle case — a customer-facing search built on database pattern matching, which produces the experience of a search that does not work and is blamed on the data rather than the tool.
9. Caching
Managed in-memory caching is well known and under-applied, largely because the caching is done at only one layer.
The layers available, all of which are worth considering: the content delivery network for static and cacheable dynamic responses; an application-level cache for computed results and session data; a database query cache for expensive repeated queries; and an API layer cache for responses that do not change per request.
The discipline that makes caching safe rather than dangerous: decide the invalidation strategy before adding the cache. Time-based expiry is simple and serves stale data. Event-based invalidation is correct and requires knowing every path that changes the underlying data. Choosing deliberately, and documenting the choice, is what prevents the class of bug where the cache is right and the data is not.
10. Identity for your users
Authentication is a solved problem that teams keep re-solving, usually incompletely.
A managed user identity service provides registration, sign-in, password reset, multi-factor authentication, social and enterprise sign-in, token issuance and user management. Building the equivalent means building all of it, and the parts most often skipped — account recovery, rate limiting on authentication endpoints, session invalidation, secure token handling — are exactly the parts that produce incidents.
The honest trade-off: these services are opinionated, and customising flows beyond what they anticipate ranges from awkward to impossible. Evaluate against your actual requirements early, because discovering a constraint after building around it is expensive.
Distinct from this is workforce identity — how your own staff and services get permissions. The under-used capability here is short-lived credentials issued to workloads rather than long-lived keys stored in configuration. Every long-lived key in a repository or an environment file is a permanent exposure, and the mechanism to eliminate them exists and is not widely adopted.
11. Secrets and configuration
Two related services worth separating.
A secrets manager for credentials, with automatic rotation, fine-grained access control and audit. The rotation capability is the differentiator — a secret rotated automatically on a schedule closes a class of risk that manual processes leave open indefinitely, because manual rotation is always deferred.
A parameter store for non-secret configuration, versioned and hierarchical. Cheaper than the secrets manager and appropriate for anything that is configuration rather than credential.
The practice worth adopting: nothing sensitive in environment variables set at deploy time, and nothing at all in the repository. Fetch at runtime, with access governed by the workload's identity. This is a small change that removes the most common source of credential leaks.
12. Networking you probably under-use
Networking tends to be configured once and never revisited, which leaves several capabilities unused.
Private endpoints to services, so traffic to managed services never traverses the public internet. Better security posture, and frequently lower data transfer cost — a rare case where the secure option is also the cheaper one.
Private connectivity between accounts or organisations, so a service can be exposed to a partner or another part of the business without a public endpoint.
Flow logs for network traffic, which are what let you answer "what is actually talking to what" — a question most teams cannot answer about their own infrastructure.
Managed firewall and threat detection at the network layer, which catches things application-level controls do not.
Global load balancing with health-based routing, which turns a multi-region deployment from a manual failover procedure into an automatic one.
13. Content delivery and the edge
A content delivery network is usually deployed for static assets and stops there, which leaves most of its value unused.
Beyond caching files: caching API responses where they are cacheable at all, which removes load from origin infrastructure entirely; terminating connections close to users, which improves latency even for uncacheable content because the slow part of a connection is the handshake; and running code at the edge for request manipulation, authentication checks, redirects, header rewriting and simple personalisation.
Edge compute is the genuinely under-used part. Logic that runs before a request reaches your infrastructure — routing by geography, rejecting unauthenticated requests, serving a maintenance page, running an experiment split — executes closer to the user and never consumes origin capacity.
14. Observability
Most teams have logs and basic metrics. The rest is available and unconfigured.
Distributed tracing shows a request's path across services with timing at each hop. In any system with more than three services, this is the difference between finding a latency problem in minutes and finding it in days.
Structured log querying rather than text search, so logs can be aggregated and analysed rather than only read.
Synthetic monitoring — scripted checks running continuously from outside — which detects failures before users report them, and detects the class of failure where every component is healthy and the journey is broken.
Real user monitoring, which measures what users actually experience rather than what your servers report.
Anomaly detection on metrics, which catches deviations from normal patterns without requiring someone to have set a threshold in advance for a metric nobody thought about.
The pattern behind all of these: they answer questions during an incident that logs alone cannot, and configuring them beforehand is what makes an incident short.
15. Governance and guardrails
Once there is more than one account or team, governance services stop being bureaucracy and start being what prevents mistakes.
Organisation-wide policies that constrain what can be done regardless of individual permissions — no resources in unapproved regions, no disabling of logging, no public storage buckets. A guardrail is more reliable than a policy document.
Configuration compliance monitoring that continuously evaluates resources against rules and reports drift, which catches the thing someone changed manually during an incident and never changed back.
Automated security findings aggregating across services, so vulnerabilities and misconfigurations arrive in one place rather than in five consoles nobody opens.
Threat detection analysing account activity and network traffic for suspicious patterns, which is the only realistic way to notice credential misuse.
Account provisioning with a baseline, so a new account arrives with logging, guardrails and network configuration already applied rather than being configured from memory.
16. Cost control that works
In rough order of impact, and the order is not what most teams assume.
Commitment-based discounts for steady baseline usage. The largest single lever for most organisations, frequently unused because it requires committing to a spend level. Analyse the actual baseline and commit to that, not to peak.
Spot capacity for anything interruptible. Large discounts for modest engineering effort.
Storage lifecycle policies. Data that has not been accessed in ninety days rarely needs to be on the most expensive tier, and the transition is a configuration rule rather than a project.
Right-sizing. Instances provisioned for an estimated peak that never occurred. Recommendation tooling exists and is ignored.
Shutting down non-production outside working hours. Development environments running continuously for a team that works forty hours a week is roughly a seventy-five percent waste on that line.
Data transfer awareness. Cross-region and internet egress charges accumulate invisibly and are frequently a top-three cost line that nobody has looked at.
Tagging and allocation. Not a saving itself, and the precondition for every other saving — you cannot optimise spend you cannot attribute.
Budgets and anomaly alerts, so a runaway cost is noticed in a day rather than in a monthly invoice.
17. Migration and modernisation services
Worth knowing about because they are used once and forgotten between projects.
Discovery tooling that inventories an existing estate and maps dependencies — answering the question every migration starts with and few teams can answer from documentation. Database migration services that handle continuous replication so cutover is short rather than a weekend. Bulk data transfer for volumes where the network is genuinely the bottleneck. And application migration tooling that replicates servers to the cloud with a short cutover.
The common thread is that these compress the risky part of a migration — the cutover window — by keeping source and target synchronised until the moment of switching.
18. How to decide
A short evaluation that avoids both the reflexive build and the reflexive adopt.
- Is this our differentiator? If not, the default should be to adopt.
- What does self-hosting genuinely cost? Including patching, scaling, on-call and the knowledge concentrated in one person.
- Does the service's model fit our problem? Check the limits and constraints before building around it, not after.
- What does exit look like? Not whether lock-in exists, but how much code would change. A queue behind an interface is cheap to leave; a data model shaped by a specific store is not.
- What does it cost at our projected scale? Managed services are usually cheaper at low and variable volume and can be more expensive at high steady volume.
- Can we prototype it in a day? Most of these can be tried quickly, and a day of experimentation beats a week of comparison.
19. Twelve mistakes
- Building a job queue in the database. A managed queue exists and handles retries and failures properly.
- Coordinating multi-step processes in application code. The state and retry logic becomes the fragile part.
- Full-text search on a relational database. Produces a search experience users describe as broken.
- Building authentication. The parts usually skipped are the parts that cause incidents.
- Long-lived credentials in configuration. A permanent exposure with an available fix.
- No dead-letter destination. Silently discarded work.
- Loading data into a warehouse to answer one question. Query it in object storage instead.
- Raw text files in object storage for anything queried repeatedly. Columnar and partitioned costs a fraction.
- Caching without deciding invalidation. Correct cache, wrong data.
- Non-production environments running continuously. Roughly three quarters wasted.
- No cost allocation tags. Every other optimisation is blocked by it.
- Ignoring commitment discounts because nobody analysed the baseline.
20. A worked example: replacing three in-house systems
A company with a mature application on virtual machines and a managed database, reviewing what it maintains that it should not.
System one: the job runner. Background work — emails, exports, webhook deliveries, report generation — ran through a jobs table in the database with a worker process polling it. About nine hundred lines of code accumulated over four years, handling retries, backoff, priorities and stuck-job recovery. It worked, and it produced an incident roughly quarterly, usually a job that failed in a way the retry logic did not anticipate and blocked the queue behind it.
Replaced with managed queues — three of them, split by priority, each with a dead-letter destination — consumed by serverless containers that scale to zero when idle. The nine hundred lines became about eighty. The dead-letter queues immediately revealed something the old system had hidden: roughly forty jobs a week were failing permanently and being silently dropped, mostly webhook deliveries to endpoints that no longer existed. Nobody had known, because the old table had no distinction between "will retry" and "gave up".
System two: the order processing flow. An eleven-step process — validate, reserve stock, charge, generate documents, notify, schedule dispatch, with branches for pre-orders and a manual approval step for high-value orders. Implemented as a service holding state in the database with a scheduler advancing stuck orders. The most feared part of the codebase, because a failure mid-flow left an order in an ambiguous state that required manual database repair.
Replaced with a workflow service. Each step became a discrete function; the transitions, retries, timeouts and the human approval wait became configuration. The immediate benefit was not the code reduction, which was moderate, but the visibility: a failed order now shows exactly which step failed, what the input was, what the error said, and can be resumed from that point. Manual database repair stopped entirely.
System three: log analysis. Application logs were shipped to object storage and, when someone needed to investigate, downloaded and grepped locally. For anything spanning more than a day this was impractical, so most investigations simply did not happen.
Replaced by converting the log delivery to a columnar format partitioned by date and querying object storage directly. No warehouse, no pipeline, no infrastructure. A question that previously took an afternoon of downloading and grepping now takes a query that scans a few gigabytes and costs a fraction of a penny. The format conversion and partitioning were the whole of the work, and they reduced query cost by roughly a factor of twenty compared to querying the raw text.
What was deliberately not replaced. The team evaluated a managed identity service and declined, because the application has an unusual multi-tenant permission model that would have required fighting the service's assumptions. That is the correct outcome of an evaluation, and it was reached in two days of prototyping rather than after three months of building around a constraint.
What it cost and saved. The direct infrastructure bill rose slightly. The maintenance burden dropped substantially — three systems requiring understanding became configuration, the quarterly job-queue incident stopped occurring, and the forty silently failing webhooks a week became a visible list somebody fixed. The engineering time recovered exceeded the additional spend by a wide margin, which is the calculation that matters and the one that does not appear on an invoice.
21. Frequently asked questions
How do I decide between building and adopting?
Start by asking whether it is your differentiator — if not, default to adopting. Then price self-hosting honestly, including patching, scaling, on-call and the fact that the knowledge usually lives in one person's head. That total is consistently underestimated because it never appears on an invoice, which is why managed services look more expensive than they are.
Isn't lock-in a serious risk?
It is proportional to how deeply the service shapes your code, not to whether you use a managed service at all. A queue behind an interface is cheap to replace. A data model designed around a particular store's access patterns is not. Evaluate exit cost per service rather than treating lock-in as a single blanket concern.
What is the highest-value thing most teams are not using?
Querying object storage directly for data already sitting there, and workflow services for multi-step processes currently coordinated by application code. Both replace substantial in-house complexity with configuration, and both are usually a day of work to prototype.
Where does cloud spend actually go?
Usually compute that was sized for a peak that never happened, non-production environments running around the clock, storage on the most expensive tier long after anyone read it, and data transfer nobody has looked at. Commitment discounts on the real baseline and lifecycle policies on storage are typically the two largest available levers.
Are serverless functions right for our services?
For event handlers, glue and sporadic work, yes. For a continuously running service under steady load, serverless containers are usually a better fit — same operational simplicity, no duration limits, fewer cold-start concerns, and a cost profile that does not punish sustained traffic.
When is a specialised database worth the complexity?
When the access pattern is genuinely the thing a relational database handles poorly — extremely high throughput on simple key lookups, deep relationship traversal, or write-heavy time-series data. A managed relational database remains the right default, and the specialised stores are answers to specific problems rather than general upgrades.
What should we configure that we probably have not?
Distributed tracing, dead-letter destinations on every queue, private endpoints to managed services, storage lifecycle policies, cost allocation tags and budget alerts. Each is a modest configuration change, and each is the thing you wish had been in place during the incident where it was missing.
How do we evaluate without committing?
Prototype for a day. Nearly all of these can be tried at negligible cost, and a day of hands-on experimentation reveals the constraints that matter far faster than reading comparisons. The failure mode to avoid is building around a service before checking whether its model fits your problem.
Key takeaways
- Price self-hosting honestly. The operational cost is real and invisible on the invoice.
- Workflow services replace the most fragile code in most systems — multi-step coordination.
- Query data where it sits. Columnar and partitioned, in object storage, with no pipeline.
- Dead-letter destinations, always. Silent work loss is the default without them.
- Commitment discounts and lifecycle policies are the two largest cost levers, and both are usually unused.
- Evaluate exit cost per service rather than treating lock-in as one undifferentiated risk.
The services worth adopting are rarely the exciting ones. They are the ones that replace a few hundred lines of coordination logic, a background job table, or a log-grepping ritual — the code nobody wanted to write and everyone has to maintain. Finding those is mostly a matter of noticing what your team keeps rebuilding.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch