Skip to main content
Blog

Serverless Architecture: When to Go Function-as-a-Service

Last updated Architecture

Serverless promises that you write a function, someone else runs it, and you pay only when it executes. That promise is real, and it comes with a set of constraints that are equally real and considerably less advertised: your code starts cold, holds no state, times out, and lives inside an event-driven architecture whether or not you intended to build one. Serverless is not simpler infrastructure. It is different infrastructure, with a different set of things to think about.

This guide covers when function-as-a-service genuinely fits, how the execution model shapes design, what cold starts actually cost, where the economics turn against you, and how to build systems that stay debuggable when they are made of a hundred small pieces.


What you will learn
  • What "serverless" covers, and the constraints that come with it
  • The execution model, and why it explains almost every design rule
  • Cold starts: the real numbers and the mitigations that work
  • Event-driven patterns, and the reliability properties they require
  • The cost curve, and the point at which containers become cheaper
  • Observability, testing and the operational reality
In this article
  1. What serverless actually means
  2. The execution model
  3. Cold starts
  4. Statelessness and its consequences
  5. Event-driven design
  6. Reliability: retries, ordering and idempotency
  7. Composing functions
  8. The database problem
  9. The cost curve
  10. Observability
  11. Testing and local development
  12. Security
  13. Where serverless fits
  14. Where it does not
  15. Twelve mistakes
  16. A worked example: an image processing pipeline
  17. Frequently asked questions

1. What serverless actually means

The term covers a spectrum, and being precise about which part you mean prevents most confused conversations.

KindExampleWhat you give up
FunctionsCode triggered by an event, billed per invocationLong execution, warm state, control over runtime
Serverless containersA container image scaled to zero automaticallyLess than functions; a middle ground
Managed data servicesDatabases and queues with no capacity to provisionFine-grained tuning
Edge functionsSmall code running at delivery pointsExecution time, memory, available libraries

The unifying property is not the absence of servers — there are servers — but the absence of capacity management. You do not decide how many instances exist. That single change is what produces both the benefits and the constraints.

This article focuses on functions, because that is where the design implications are strongest and where the decision is least obvious.

2. The execution model

Almost every rule below follows from how functions actually execute, so it is worth stating precisely.

An event arrives. The platform finds an available execution environment or creates one. Your code runs. The environment is kept warm for a while in case another event arrives, then discarded. Concurrency is achieved by running many environments in parallel, each handling exactly one event at a time.

Four consequences shape everything:

  • One event per environment at a time. There is no in-process concurrency to exploit; a function waiting on a slow API is billed for that waiting and cannot serve another request meanwhile.
  • Environments are created and destroyed unpredictably. Anything in memory may vanish. Anything written to local disk may vanish.
  • Initialisation code runs once per environment, not per event. This is the most useful property to exploit: connections and clients created outside the handler are reused across invocations in the same environment.
  • Execution is bounded. There is a maximum duration, and exceeding it terminates the function mid-work.

The scaling behaviour is what makes functions attractive: from zero to thousands of concurrent executions in seconds, with no configuration. It is also what makes them dangerous downstream, since a burst of concurrency arrives at your database as a burst of connections.

3. Cold starts

A cold start is the additional latency when no warm environment exists: the platform must provision one, load your code, and run your initialisation before your handler executes.

The magnitude varies enormously by runtime and by what you do at initialisation. Interpreted runtimes with small dependency trees start fastest. Runtimes requiring a virtual machine and class loading start slowest. A function that opens database connections and loads configuration at initialisation adds that time to every cold start.

What actually reduces them, in order of effectiveness:

  • Ship less code. Package size directly affects load time. Bundle and tree-shake; exclude development dependencies; do not package an entire cloud SDK when you use one client from it.
  • Do less at initialisation. Create clients lazily where they are not always needed. Fetch configuration once and cache it, rather than on every cold start if the value is stable.
  • Choose the runtime deliberately when latency is user-facing. This is one of the few cases where language choice has a direct product consequence.
  • Allocate more memory. Processor allocation scales with memory on most platforms, so a larger allocation initialises faster — and frequently costs the same or less overall, because duration falls.
  • Provisioned concurrency keeps environments warm at a fixed cost. Effective, and it erodes the pay-per-use economics, so it should be a deliberate trade rather than a default.

The honest framing: cold starts matter for user-facing synchronous requests and matter very little for asynchronous background work. A queue consumer taking an extra second to start is invisible; a checkout endpoint doing the same is not.

4. Statelessness and its consequences

No state survives reliably between invocations, and designing around that is the main adjustment for teams coming from long-running services.

Session state must live elsewhere — a cache, a token, a database. Anything held in a module-level variable will be present for some requests and absent for others, which produces bugs that are intermittent and maddening to reproduce.

In-memory caching still works, partially. A value cached at module level is reused within a warm environment, which is genuinely useful — but with a hundred concurrent environments you have a hundred independent caches with independent expiry. Treat it as a per-instance optimisation, never as a shared cache.

Background work after responding is not reliable. Once the handler returns, the environment may be frozen immediately, and work started but not awaited may simply never complete. If something must happen, await it or hand it to a queue.

Local disk is scratch space, sized modestly and shared across invocations in the same environment. Useful for temporary files during processing; never for anything that must persist or that another invocation depends on.

5. Event-driven design

Functions are triggered by events, which means adopting them is adopting event-driven architecture whether or not that was the intention. The common sources:

TriggerTypical useNote
HTTP requestAPIs and webhooksSynchronous; cold starts are user-visible
Queue messageBackground processingRetries and dead-letter handling included
Object storage eventFile processing on uploadAt-least-once; expect duplicates
Stream recordChange data capture, analyticsOrdered within a partition; batch processing
SchedulePeriodic jobsSimplest replacement for a cron server
Event busReacting to system eventsRouting rules as configuration, not code

The design principle that matters most is one function, one responsibility. A function that handles six event types with a switch statement inherits the worst of both models: the complexity of a service without the ability to share warm state or scale as a unit. Small focused functions are easier to reason about, scale independently, and fail in isolation.

The counter-pressure is real: hundreds of tiny functions become difficult to navigate, deploy and trace. The workable middle is a function per meaningful operation rather than per line of logic, grouped and deployed together as a coherent service.

6. Reliability: retries, ordering and idempotency

Distributed event-driven systems have failure modes that a single process does not, and the platform's defaults assume you have handled them.

Delivery is at-least-once. The same event will occasionally be delivered twice — after a timeout, after a partial failure, or because the platform could not confirm your success. Every handler must therefore be idempotent, keyed on something stable in the event. This is not defensive programming; it is a requirement of the model.

Retries are automatic and configurable. Understand the specific behaviour for each trigger type, because it differs: synchronous invocations typically do not retry, asynchronous ones retry a fixed number of times, and queue-based ones retry until the message succeeds or reaches a dead-letter destination.

Dead-letter destinations are not optional. Without one, a message that always fails is retried indefinitely, consuming capacity and obscuring the real problem. With one, failures accumulate somewhere inspectable, and you can alert on the depth of that queue.

Ordering is usually not guaranteed. Stream-based triggers preserve order within a partition; queues and event buses generally do not. Designing handlers that tolerate out-of-order arrival — by fetching current state rather than inferring it from sequence — avoids an entire category of intermittent bug.

Partial batch failure deserves explicit handling. When a function processes a batch and one item fails, the default on some platforms is to retry the whole batch, reprocessing the successful items. Reporting per-item failure is a configuration option worth using.

7. Composing functions

Multi-step processes need coordination, and there are three approaches with quite different properties.

Chaining through events. Each function emits an event that triggers the next. Loosely coupled and simple to build; the overall process exists nowhere explicitly, which makes it difficult to understand, monitor or reason about when a step fails halfway.

Orchestration with a workflow service. The sequence is defined as a state machine with retries, branching, error handling, parallel steps and long waits handled by the platform. The process is visible and its history is inspectable. This is the right answer for anything with more than about three steps, anything with compensation logic, or anything running longer than a function's maximum duration.

A single longer function. Sometimes correct, and frequently dismissed too quickly. If three steps always run together, take a few seconds, and share data, one function is simpler than three plus the plumbing between them.

The mistake to avoid is a chain of six functions connected by events with no orchestration, where a failure at step four leaves the system in a state nobody can describe and no interface can show.

8. The database problem

This is the most common serious problem in production serverless systems, and it is structural rather than accidental.

Traditional databases handle a modest number of persistent connections from a small number of long-lived application instances. Functions scale to hundreds of concurrent environments, each wanting its own connection. The database exhausts its connection limit, and every function starts failing — including the ones that were working perfectly a moment earlier.

The mitigations, in order of preference:

  • A connection pooler between functions and the database, which multiplexes many function connections onto few database connections. This is the standard answer for relational databases and should be assumed necessary rather than added after the incident.
  • Data services designed for this model — key-value stores accessed over HTTP with no connection concept — which sidestep the problem entirely.
  • Reserved concurrency capping how many instances of a function may run at once, which protects the database at the cost of queueing work.
  • Connection reuse at module level, so a warm environment reuses its connection rather than opening one per invocation. Necessary but insufficient on its own.

The broader lesson generalises beyond databases: anything downstream of a function must tolerate the function's scaling behaviour. A third-party API with a rate limit will be hit at full concurrency the first time traffic spikes.

9. The cost curve

Serverless billing is per invocation and per unit of duration multiplied by allocated memory. That produces a distinctive economic shape.

Extremely cheap at low and spiky volume. A function invoked a few thousand times a day costs almost nothing, and a service used heavily for four hours daily costs a fraction of a continuously running container. This is the strongest economic argument and it is genuine.

More expensive at sustained high volume. Beyond a certain steady request rate, a continuously running container serving the same traffic is cheaper — often substantially. The crossover depends on execution duration and memory, but it exists and it is worth calculating rather than assuming.

Three cost traps worth naming:

  • Paying for waiting. A function idle on a slow third-party call is billed throughout. Long synchronous waits are the least economical thing you can do in this model.
  • Under-allocated memory. Because processing power scales with memory, a smaller allocation can run long enough that it costs more than a larger one. Memory should be tuned empirically, and the optimum is frequently higher than intuition suggests.
  • Recursive triggers. A function writing to the storage bucket that triggers it creates an infinite loop billed per invocation. This is a well-known way to generate a memorable invoice, and it is why concurrency limits and budget alerts belong in place from day one.

10. Observability

Debugging is genuinely harder here, because there is no server to inspect, no persistent process, and a request may traverse a dozen functions and queues.

Structured logging with a correlation identifier propagated through every event payload. Without it, reconstructing a single logical request from logs across many functions is close to impossible.

Distributed tracing is not optional at any meaningful scale. It is the only practical way to answer where the time went and which step failed.

Metrics that matter for functions specifically: invocation count, error rate, duration distribution, throttles, concurrent executions, and cold start rate. Duration averages hide the story — the ninety-fifth percentile is where cold starts live.

Alert on the things that indicate real problems: dead-letter queue depth above zero, throttling, error rate, and duration approaching the timeout. That last one is an early warning that something is degrading before it starts failing outright.

11. Testing and local development

The local development story is weaker than for containers, and pretending otherwise leads to frustration.

The approach that works: keep business logic in plain functions with no dependency on the platform's event shapes, and make the handler a thin adapter that translates an event into a call to that logic. The logic is then testable as ordinary code with ordinary unit tests, which is fast and reliable.

For integration testing, local emulators exist and approximate the platform imperfectly — they are useful for the shape of an event and misleading about permissions, timeouts and concurrency. The higher-fidelity approach is deploying to a real ephemeral environment per branch, which is practical because infrastructure as code makes environments cheap.

The tests worth writing explicitly, because they cover what breaks in production: duplicate event delivery, out-of-order arrival, a downstream dependency timing out, and a batch where one item fails.

12. Security

  • One role per function, with exactly the permissions that function needs. The convenience of a shared permissive role is how a minor vulnerability in one function becomes access to everything.
  • Secrets from a secrets manager, fetched at initialisation and cached in the environment — not embedded in environment variables, which appear in configuration and deployment output.
  • Validate every event. An event from a queue or bus is input from another system and may be malformed or malicious. The absence of an HTTP boundary does not make it trusted.
  • Dependency scanning in the pipeline, because functions frequently accumulate large dependency trees for small amounts of code.
  • Concurrency limits as a blast radius control — for cost, for downstream protection, and to bound the damage from a runaway trigger.

13. Where serverless fits

The workloads where it is clearly the right answer share recognisable properties.

  • Spiky or unpredictable traffic. Paying nothing at idle and scaling instantly is exactly what the model provides.
  • Event-driven background processing. File uploads, webhooks, notifications, data transformation — where a cold start is invisible and each unit of work is independent.
  • Scheduled jobs. The simplest possible replacement for a machine that exists to run cron.
  • Glue between services. Reacting to an event, transforming a payload, calling an API. Too small to justify a service.
  • Low-volume internal tools. Where a running container would cost more than the tool is worth.
  • Genuinely bursty batch work. Fanning out to hundreds of parallel executions for a few minutes and back to zero.

14. Where it does not

  • Sustained high-volume services. The economics reverse, and containers become both cheaper and more predictable.
  • Strict low-latency requirements. Cold starts introduce variance that provisioned concurrency mitigates at a cost that undermines the model.
  • Long-running work. Anything beyond the maximum duration needs a different execution model or an orchestrated sequence of steps.
  • Persistent connections. Websockets and streaming require platform-specific arrangements that add complexity.
  • Heavy shared state or large in-memory caches. Per-environment caching is a poor substitute.
  • Applications requiring specific runtimes or system libraries, where packaging becomes an ongoing burden.

The pattern in both lists: serverless suits work that is short, independent, and variable in volume. It suits poorly work that is long, stateful, or steady.

15. Twelve mistakes

  1. Non-idempotent handlers. At-least-once delivery means duplicates are certain.
  2. No dead-letter destination. Poison messages retried forever, obscuring the real failure.
  3. Direct database connections without pooling. Connection exhaustion under the first real spike.
  4. Work started after the response. The environment may freeze immediately; the work never completes.
  5. Module-level state treated as shared. Present for some requests, absent for others.
  6. Long synchronous waits. Paying for idle time is the least economical use of the model.
  7. Memory set to the minimum. Frequently costs more, because duration rises faster than memory price falls.
  8. No concurrency limits. A runaway trigger becomes a memorable invoice.
  9. A function writing to the bucket that triggers it. The classic infinite loop.
  10. Chains of functions with no orchestration. A failure at step four leaves a state nobody can describe.
  11. Business logic in the handler. Untestable without the platform.
  12. Adopting it for sustained high-volume workloads. Paying a premium for elasticity you never use.

16. A worked example: an image processing pipeline

Consider a requirement that fits the model well: users upload images, which must be validated, resized into several variants, scanned for inappropriate content, and recorded against the user's account. Volume is spiky — quiet overnight, heavy during business hours, occasional bursts when a customer bulk-uploads a catalogue.

The upload does not go through a function. The application requests a signed upload location, and the browser uploads directly to object storage. Passing a large binary through a function would mean paying for the transfer duration and hitting payload limits for no benefit. This is the first design decision and the one that most often goes wrong.

Storage emits an event, which enqueues rather than invoking directly. Triggering the processing function straight from the storage event works and gives you no control over concurrency. Putting a queue in between means a bulk upload of two thousand images does not open two thousand concurrent database connections — the queue absorbs the burst and the function's reserved concurrency drains it at a rate the rest of the system tolerates.

Processing is idempotent by construction. The output variant is written to a deterministic key derived from the source object and the variant name. A duplicate delivery overwrites an identical file and updates the same record, which costs a little work and produces no incorrect state. This matters because storage events are at-least-once, and duplicates are not rare.

The function does its work with generous memory. Image resizing is processor-bound, and processing power scales with memory allocation, so a larger allocation completes faster. Measured empirically, the higher allocation costs less per image than the minimum did, because duration falls faster than the price rises — a result that surprises people and is common for compute-bound functions.

Content scanning is a separate step, orchestrated rather than chained. The sequence — validate, resize, scan, record, notify — is defined as a workflow rather than as five functions emitting events at each other. The result is that a failure at the scanning step is visible as a stopped execution at a named state, with its input inspectable and a retry available, rather than as an image that silently never appeared.

Failures land in a dead-letter queue with an alert on depth. The first week surfaces two genuine categories: images in an unexpected format, and one file large enough to exceed the function timeout. Both are fixed — one by validation, one by routing large files to a longer-running container job — and neither would have been noticed without the dead-letter queue.

The cost profile validates the choice. Overnight the system costs nothing. The business-hours load costs a small fraction of what a continuously provisioned worker pool would. And the occasional two-thousand-image burst completes in minutes rather than hours, because concurrency scales to meet it without anyone provisioning anything.

Change one property — make it a continuous stream of images at a steady high rate all day — and the calculation reverses, with a container-based worker pool becoming both cheaper and more predictable. That is the decision in one sentence: variability is what serverless is worth paying for.

17. Frequently asked questions

Is serverless cheaper than containers?

At low, spiky or intermittent volume, dramatically. At sustained high volume, usually not — a continuously running container serving the same traffic is frequently cheaper. Calculate it for your actual profile: invocations, average duration and memory, against the cost of instances that would serve the same load. The crossover is real and worth knowing rather than assuming.

How bad are cold starts really?

It depends entirely on runtime, package size and initialisation work, and on whether the latency is user-facing. For asynchronous background work they are irrelevant. For a synchronous API endpoint they add variance that shows up at the ninety-fifth percentile. Reduce package size and initialisation work first; reach for provisioned concurrency only when the latency requirement genuinely demands it.

Can we build a whole application this way?

Yes, and many organisations have. It works best when the application is genuinely event-driven and traffic is variable. It works less well for a conventional request-response application with steady traffic, where you inherit cold starts, connection management and distributed debugging in exchange for elasticity you are not using. A hybrid — functions for events and background work, containers for the steady API — is a very common and sensible outcome.

How do we handle relational databases?

A connection pooler between functions and the database, treated as a requirement rather than an optimisation. Reuse connections at module level so warm environments do not reconnect, and set reserved concurrency to bound the load. Where the access pattern allows, a data service accessed over HTTP with no connection concept removes the problem entirely.

What about vendor lock-in?

It is real, and mostly in the event integrations and platform services rather than in the function code itself. The practical mitigation is keeping business logic in plain functions with the handler as a thin adapter — that logic moves easily. The surrounding architecture of triggers, queues and workflows is where the migration cost sits, and it is genuine. Whether that matters depends on how likely a move actually is.

How do we test properly?

Unit test the business logic as ordinary code, because it has no platform dependency if you structured it correctly. Use ephemeral deployed environments for integration testing rather than relying on local emulators, which approximate permissions and concurrency poorly. Explicitly test duplicate delivery, out-of-order arrival, downstream timeouts and partial batch failure — those are what break in production.

How many functions should a service have?

One per meaningful operation, grouped and deployed together as a coherent unit. A single function handling many event types with a switch statement loses the isolation benefit; hundreds of trivial functions become impossible to navigate and trace. The useful test is whether each function has a name that describes one thing someone would ask for.

What is the first thing to get right?

Idempotency, and a dead-letter destination on every asynchronous trigger. Duplicates and permanent failures are certainties in this model, not edge cases, and both produce problems that are difficult to diagnose after the fact. Concurrency limits and a budget alert are a close second, because the failure mode there is financial rather than technical.

Key takeaways

  • Serverless removes capacity management, not complexity. It trades one set of concerns for another.
  • Variability is what you are paying for. Spiky and intermittent workloads win; steady high volume does not.
  • Idempotency and dead-letter destinations are requirements, because delivery is at-least-once and failures are permanent.
  • The database is the usual failure point. Pool connections and bound concurrency before the first real spike.
  • Keep logic out of handlers. It is what makes the code testable and portable.
  • Orchestrate anything with more than three steps. Chains of event-triggered functions become unexplainable.

The best serverless systems are unremarkable: small independent functions doing short pieces of work, triggered by events, idempotent by construction, with failures landing somewhere visible. Reaching for it because it is fashionable produces distributed complexity; reaching for it because your traffic is genuinely variable produces a system that costs nothing when nobody is using it.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch