Skip to main content
Blog

IoT Architecture Explained: From Edge Devices to Cloud Dashboards

Last updated Architecture

A temperature reading leaves a sensor in a factory and, ninety seconds later, appears on a dashboard three thousand kilometres away. Between those two moments it passes through six or seven distinct systems, each with its own failure modes, and any one of them can silently drop it without anybody noticing for a month.

This article maps that path end to end — devices, connectivity, gateways, protocols, ingestion, storage, processing and presentation — and pays particular attention to the parts that look simple in a diagram and consume most of the project in practice: provisioning, updates, security and the sheer awkwardness of physical hardware.


What you will learn
  • The layers of an IoT system and what each is responsible for
  • Choosing connectivity, and why it constrains everything else
  • What gateways do and when you need one
  • Protocol choices and their real trade-offs
  • Storing and querying time-series data at volume
  • Device management, updates and security
In this article
  1. The shape of an IoT system
  2. The device layer
  3. Power, and why it decides the design
  4. Connectivity options
  5. Gateways and edge computing
  6. Protocols
  7. Identity and provisioning
  8. Ingestion
  9. Storing time-series data
  10. Processing: stream and batch
  11. The digital twin
  12. Dashboards and alerting
  13. Device management and updates
  14. Security across the stack
  15. Scaling, and where it breaks
  16. Cost drivers
  17. Twelve mistakes
  18. A worked example: cold chain monitoring
  19. Frequently asked questions

1. The shape of an IoT system

Almost every deployment has the same layers, whatever the domain.

LayerResponsibilityTypical concerns
DeviceSense, actuate, bufferPower, cost, reliability, physical environment
ConnectivityMove data off the deviceRange, bandwidth, power, coverage, cost per device
Gateway / edgeAggregate, translate, filter, act locallyLocal processing, protocol bridging, buffering
IngestionAccept and authenticate at volumeThroughput, backpressure, device identity
StorageRetain raw and derived dataVolume, retention, query patterns, cost
ProcessingTransform, aggregate, detectLatency, correctness, late data
PresentationDashboards, alerts, integrationsUsability, alert fatigue, access control
ManagementProvision, configure, update, monitorFleet visibility, safe rollout, recovery

The management layer is drawn last and matters most. A system with a thousand devices spends far more effort on keeping them registered, configured, updated and healthy than on the data path itself — and teams that treat management as an afterthought discover this eighteen months in, when a third of the fleet is running firmware nobody can identify.

2. The device layer

Devices range from a battery sensor costing a few pounds to an industrial controller costing thousands. The architecture differs enormously across that range, but the responsibilities are consistent: read sensors, apply local logic, buffer when the network is unavailable, and transmit.

Buffering deserves emphasis because it is routinely omitted. Networks fail. A device that discards readings when it cannot connect produces gaps that are invisible in the data and destroy any analysis depending on continuity. Even a modest local buffer — an hour, a day — converts a data loss into a delay, and that distinction matters enormously downstream.

Local logic reduces everything else. A sensor sampling every second and transmitting every reading generates eighty-six thousand messages a day per device. The same sensor transmitting only on meaningful change, plus a heartbeat, might send a few hundred. That is a hundred-fold reduction in bandwidth, ingestion cost, storage and battery drain, achieved with a threshold comparison on the device.

The counter-consideration: filtered data cannot be un-filtered. If you later want the full-resolution signal to investigate an anomaly, it never existed. The usual compromise is aggressive filtering for transmission with a short full-resolution buffer on the device that can be requested on demand.

3. Power, and why it decides the design

For battery devices, power is the constraint everything else bends around, and it is worth understanding why.

Radio transmission dominates energy consumption — typically by an order of magnitude over sensing and computation. A device that transmits once an hour may last years; the same device transmitting once a minute may last weeks. This single fact determines sampling rates, protocol choice, connectivity technology and the entire data model.

The consequences ripple outward. Devices sleep most of the time, which means they cannot be reached on demand — commands must queue until the device wakes. Protocols with persistent connections are unsuitable because maintaining a connection costs power. Firmware updates are expensive because transferring an image is a large transmission, so update strategies favour small differential patches.

Where mains power is available, most of this relaxes and a far simpler architecture becomes possible. Establishing early which devices are battery-powered and which are not is one of the highest-leverage decisions in the project, because it partitions the design space.

4. Connectivity options

TechnologyRangeBandwidthPowerSuits
Wi-FiTens of metresHighHighMains-powered devices in covered buildings
Bluetooth low energyTens of metresLowVery lowWearables, devices paired to a nearby gateway
Cellular (LTE-M, NB-IoT)Wide areaLow to moderateModerateMobile or dispersed assets with existing coverage
LoRaWANKilometresVery lowVery lowSparse sensors over a site or region
Zigbee / ThreadTens of metres, meshedLowLowDense indoor deployments
Wired (Ethernet, industrial buses)FixedHighNot applicableFixed industrial equipment

Three practical observations that matter more than the table.

Coverage is a site question, not a specification question. A technology's stated range assumes conditions your site does not have. Concrete, metal, machinery and geography all interfere. Survey before committing, because discovering coverage gaps after deploying two hundred devices is expensive in a way nothing else on this list is.

Cellular shifts cost from capital to operating. No gateway infrastructure to install, a recurring charge per device forever. For a few dozen devices that is usually cheaper; for a few thousand it usually is not.

Low-power wide-area technologies are genuinely low bandwidth. LoRaWAN messages are measured in tens of bytes with strict transmission duty limits. This is not a device you send an image from, and it constrains the data model considerably — teams that design a message format before checking the payload limit routinely have to start again.

5. Gateways and edge computing

A gateway sits between devices and the network, and does more than forward.

Protocol translation. Devices speak a local protocol; the cloud speaks another. Industrial equipment in particular speaks decades-old protocols that will never speak anything else, and the gateway is where that is bridged.

Aggregation. Fifty devices reporting to one gateway that sends one combined message reduces connection overhead dramatically.

Buffering. A gateway with storage and power can hold days of data through an outage, which is far more than a constrained device can.

Local processing. The genuinely important one. Some decisions cannot wait for a round trip: a safety interlock, a control loop, a quality check on a production line. Anything requiring a response in milliseconds must be decided locally, because network latency alone rules out the alternative.

Local operation. A well-designed edge system keeps working when the connection fails, syncing when it returns. For anything operationally significant, this is not optional — a factory that stops because a broadband link dropped is not an acceptable design.

The trade-off is that gateways are computers in the field. They need updating, monitoring, securing and occasionally physically visiting. Each one is infrastructure, and a deployment with two hundred gateways has two hundred small servers in places without staff.

6. Protocols

MQTT is the default for good reasons. It is publish-subscribe, so devices publish to topics without knowing who consumes them. Overhead is minimal. It supports quality-of-service levels — fire-and-forget, at-least-once, exactly-once — and a last-will message that the broker publishes if a device disconnects unexpectedly, which is how fleet health monitoring is usually built. Its weaknesses are that it needs a persistent connection, which costs power, and that topic hierarchies become unmanageable if not designed deliberately from the start.

CoAP is designed for the most constrained devices — request-response over a lightweight transport, no persistent connection, very small overhead. It suits battery devices that wake, transmit and sleep. Less ecosystem support than MQTT.

HTTP is heavy for constrained devices and entirely reasonable for mains-powered ones with decent bandwidth. Its advantage is that everything already speaks it, and that is worth more than protocol efficiency in many deployments.

Industrial protocols — the various fieldbus and automation standards — exist in enormous installed bases and are not going away. In practice they terminate at a gateway that translates.

The choice is less important than consistency. A deployment with four protocols has four sets of failure modes, four sets of tooling and four things to secure.

7. Identity and provisioning

Every device needs an identity the platform trusts, and getting devices from a factory into a running fleet is one of the genuinely hard operational problems.

The essential requirement: each device has a unique credential, ideally a certificate, ideally created during manufacture and stored where it cannot be extracted. Shared credentials across a fleet mean one compromised device compromises all of them, and this has happened often enough to be a well-documented pattern rather than a hypothetical.

The provisioning flow that works at scale: the device is manufactured with a credential; it is registered in the platform against an expected identity; on first connection it authenticates and receives its operational configuration; and it appears in the fleet with a known state. Anything requiring a person to type a serial number does not scale past a few hundred devices, and anything requiring a technician to configure a device on site does not scale at all.

Also plan for the reverse: decommissioning. Devices are retired, stolen, sold or replaced. A credential that is never revoked is a permanent hole, and fleets that have run for years without a revocation process invariably have hundreds of them.

8. Ingestion

The front door, and the layer where volume first becomes real.

Its responsibilities are narrow and demanding: authenticate every device, accept messages at whatever rate the fleet produces them, and hand them to processing without losing any. It should do as little else as possible — validation, enrichment and business logic belong downstream, because anything in the ingestion path is a potential source of backpressure.

The characteristic difficulty is that IoT traffic is bursty in correlated ways. Devices configured identically report on the same schedule, so a fleet of ten thousand devices reporting every five minutes produces ten thousand messages in a narrow window and near-silence between. Worse, a network outage followed by recovery produces every device reconnecting and flushing its buffer simultaneously — a burst far larger than normal peak, at exactly the moment the system is least healthy.

Two mitigations, both worth designing in from the start: jitter the reporting schedule so devices spread across the interval rather than aligning, and back off randomly on reconnection so recovery is staggered. Both are a few lines of device firmware and prevent an entire category of outage.

Behind ingestion, a durable queue decouples arrival rate from processing rate. Without it, a slow consumer becomes a lost message.

9. Storing time-series data

IoT data is overwhelmingly time-series: a device, a measurement, a timestamp, a value. This shape has specific properties that general-purpose databases handle poorly.

The properties: writes are heavy and almost always appends; data is rarely updated; queries are almost always bounded by time and device; and older data is queried less but rarely deleted.

The pattern that works is tiered retention:

  • Raw, recent. Full resolution for days or weeks. Used for investigation and detailed dashboards.
  • Downsampled, medium-term. Hourly or daily aggregates for months to years. This is what most dashboards actually query.
  • Archived. Raw data in cheap object storage, retrievable slowly if ever needed.

Storage cost is dominated by raw retention, so the downsampling policy is effectively the budget. The mistake worth avoiding is keeping everything at full resolution indefinitely because the decision felt hard — that is a decision, and it is usually the most expensive available one.

A second consideration: late and out-of-order data is normal, not exceptional. A device that buffered through an outage delivers yesterday's readings today. Aggregates computed before those arrive are wrong, and a system that cannot recompute them will show numbers that quietly disagree with the raw data.

10. Processing: stream and batch

Two paths, both usually necessary.

Stream processing handles each message as it arrives: validate it, convert units, enrich it with device metadata, evaluate alert rules, update current state. Latency is seconds. This is what makes a dashboard live and an alert timely.

Batch processing runs periodically over accumulated data: daily aggregates, trend analysis, model training, reports. Latency is hours, and correctness is easier because all the data has arrived.

The awkward middle is where most complexity lives. An alert must fire in seconds but should not fire on a single spurious reading, which means holding state across a window. A daily total should be visible during the day, which means an approximate streaming figure that a batch job later corrects. Deciding deliberately which numbers are approximate-and-live versus exact-and-delayed — and telling users which is which — prevents a great deal of confusion later.

Enrichment matters more than it sounds. A raw reading is a device identifier and a number. Useful analysis needs the device's location, its type, what asset it is attached to, and its calibration. That context lives in a device registry, and joining it in early is what turns telemetry into information.

11. The digital twin

A useful pattern: maintain a server-side representation of each device's current and desired state.

Reported state is what the device last said about itself — its readings, firmware version, configuration and health. Desired state is what it should be — the configuration and firmware you want it running.

The platform's job is reconciling them. When a device connects, it compares and applies any difference. This solves the sleeping-device problem elegantly: you set the desired state at any time, and it takes effect when the device next wakes, without anything needing to be online simultaneously.

It also gives applications something to query. Rather than asking a battery device for its current reading — which may take hours — an application reads the last reported state instantly. Almost every mature platform converges on this pattern, and building it early avoids retrofitting it later.

12. Dashboards and alerting

The layer users judge the system by, and the one most often built last and worst.

For dashboards: query the downsampled tier, not the raw one, or they will be slow and expensive. Show data freshness explicitly — a chart that silently stops updating when a device dies is worse than an empty one. And design for the fleet view as well as the device view, because with a thousand devices nobody looks at them individually.

For alerting, alert fatigue is the failure mode that kills deployments. A system that pages twenty times a day gets ignored within a fortnight, after which it provides nothing. What helps:

  • Alert on sustained conditions, not instantaneous readings. Five minutes above threshold, not one sample.
  • Distinguish device failure from condition breach. A sensor that stopped reporting is a different problem from a freezer that is too warm, and it goes to different people.
  • Suppress correlated alerts. One gateway failure should produce one alert, not fifty.
  • Every alert needs an action. If nobody knows what to do about it, it should be a dashboard metric instead.

13. Device management and updates

The part that determines whether a deployment is sustainable, and the part most likely to be underestimated.

Fleet visibility. Which devices exist, where they are, what firmware they run, when each last reported, and which are unhealthy. Without this you are operating blind, and it becomes noticeable at about the two-hundred-device mark.

Configuration management. Changing sampling rates, thresholds or endpoints across a fleet without visiting anything. Via the desired-state mechanism, and rolled out gradually rather than to everything at once.

Firmware updates over the air. Non-negotiable for anything with a multi-year life, because security vulnerabilities will be found. The requirements are strict: images must be signed and verified before installation; updates must be atomic with a fallback to the previous version if the new one fails to boot; rollout must be staged so a bad update does not brick the whole fleet; and the process must tolerate power loss mid-update.

That last point causes the most field failures. A device that loses power halfway through writing firmware and cannot boot is a device someone must physically visit — which for a sensor on a pole in a remote location may cost more than the device.

Remote diagnostics. The ability to retrieve logs and detailed state from a device without visiting it. The difference between a fifteen-minute investigation and a site visit.

14. Security across the stack

IoT security is harder than ordinary application security because devices are physically accessible, long-lived, resource-constrained and deployed in enormous numbers.

On the device: secure boot so only signed firmware runs; keys in hardware where they cannot be read out; no debug interfaces left enabled in production; and absolutely no default or shared credentials. Physical access to one device should not compromise anything beyond that device.

In transit: encrypted always, with mutual authentication so the device verifies the platform and the platform verifies the device. One-way authentication invites impersonation of the server.

At the platform: each device authorised only for its own topics and data. A compromised device that can read the entire fleet's telemetry, or publish as another device, turns one compromise into a fleet compromise.

Over the lifecycle: credential rotation, revocation on decommissioning, and a plan for what happens when a vulnerability is found in a component. Devices deployed for ten years will need patching, and a fleet without update capability is a fleet with permanent vulnerabilities.

The general principle: assume individual devices will be compromised and design so the consequence is bounded. Physical access to hardware in the field is difficult to prevent and reasonable to expect.

15. Scaling, and where it breaks

Systems tend to break at predictable points, and knowing them lets you build for the next one rather than the current one.

Around a hundred devices, manual processes stop working. Provisioning by hand, tracking devices in a spreadsheet, updating individually — all fine at ten, painful at a hundred, impossible at a thousand.

Around a thousand, correlated bursts become a real problem, storage growth becomes a real cost, and the absence of fleet-wide visibility becomes acute.

Around ten thousand, per-device operations that were cheap become expensive, alert volume overwhelms any manual handling, and anything not automated is effectively broken.

The practical advice: build the management layer for ten times your current fleet and the data layer for a hundred times. Management is hard to retrofit under load; data volume is more tractable if partitioning and retention were considered early.

16. Cost drivers

Four dominate, and their relative weight shifts with scale.

Connectivity, particularly cellular, is a recurring per-device charge that scales linearly forever. At large fleets it frequently exceeds every other running cost combined.

Ingestion is usually priced per message, which makes message frequency a direct cost lever. Halving reporting frequency halves this line.

Storage grows without bound unless retention is managed, and raw data dominates. This is the cost that surprises teams in year two.

Field operations — visiting devices — is the cost most often omitted from planning and frequently the largest. A single site visit can exceed the device's cost several times over, which is the entire economic argument for remote diagnostics and reliable updates.

17. Twelve mistakes

  1. No device buffering. Network outages become permanent data gaps.
  2. Synchronised reporting. Correlated bursts that get worse after an outage.
  3. Shared credentials. One compromise becomes a fleet compromise.
  4. No update mechanism. Permanent vulnerabilities in devices with ten-year lives.
  5. Non-atomic updates. A power cut mid-update means a site visit.
  6. Keeping everything raw forever. A storage bill that grows without a decision behind it.
  7. Ignoring late data. Aggregates that quietly disagree with the raw records.
  8. Alerting on single readings. Alert fatigue within a fortnight.
  9. No fleet view. Operating blind past a couple of hundred devices.
  10. Assuming connectivity from datasheets. Coverage is a site property; survey it.
  11. Requiring the cloud for local decisions. Safety and control loops must work offline.
  12. Omitting field operations from the budget. Frequently the largest running cost.

18. A worked example: cold chain monitoring

A distributor monitoring temperature across four hundred refrigerated units at sixty sites, plus vehicles in transit. Regulatory requirement to prove temperature was maintained; commercial requirement to catch failures before stock is lost.

Devices. Battery-powered temperature and door sensors, one per unit, with a five-year target life. That target immediately fixes several decisions: transmission must be infrequent, the protocol must be lightweight, and the device must sleep aggressively. Sampling every minute, transmitting every fifteen — or immediately if a threshold is crossed. Twenty-four hours of local buffering, which is the difference between a network outage being a delay and being a compliance failure.

Connectivity. Two answers, because the estate has two situations. Fixed sites use LoRaWAN with one gateway per site, chosen after a survey found Wi-Fi coverage unreliable inside walk-in freezers — metal enclosures being exactly the environment radio specifications describe least accurately. Vehicles use cellular, because there is no gateway to connect to and coverage gaps are handled by the device buffer.

Gateways. One per site, mains-powered, with a battery backup that matters because a power failure is precisely when temperature monitoring is most needed. Each gateway buffers seven days and runs a local rule: if any unit exceeds threshold for ten minutes, sound a local alarm regardless of whether the cloud is reachable. That rule is the reason the system is worth having, and it deliberately does not depend on the network.

Ingestion and storage. Messages arrive over MQTT with per-device certificates. Raw readings retained at full resolution for thirty days, hourly aggregates for seven years to satisfy the retention obligation, and raw data archived to object storage for the same period at a fraction of the cost. That tiering was calculated before deployment: keeping four hundred devices' minute-resolution data hot for seven years would have cost more than the hardware.

Processing. Stream processing enriches each reading with the unit's location, type and calibration offset, then evaluates alert rules. Alerts fire on ten minutes sustained above threshold, not on a single reading — a decision made after a pilot produced forty alerts a day, nearly all from door openings. Separately, a missing-data rule fires if a device has not reported in ninety minutes, routed to facilities rather than to operations, because a silent sensor is a maintenance problem rather than a stock problem.

Digital twin. Each unit has a reported state — last reading, battery level, firmware version, signal quality — and a desired state holding its thresholds and reporting interval. Changing a threshold across all frozen units is one update to desired state; the devices pick it up when they next wake. Before this existed, the same change required a technician at each site.

Dashboards. A fleet view showing sites, units within them, and any unit outside range or silent. A per-unit view with the temperature trace and door events overlaid, which is what makes an excursion interpretable — most excursions turn out to be a door propped open during a delivery. And a compliance report generated from the aggregate tier, which is the artefact the regulator actually asks for.

What went wrong in the first year. Three things, all instructive. Devices were initially configured to report on the same schedule, so every fifteen minutes four hundred messages arrived in a two-second window and the gateway occasionally dropped some; adding jitter to the reporting interval fixed it entirely. Battery life came in at three years rather than five, because the threshold-crossing transmissions were far more frequent than modelled — freezer compressors cycle, and readings hovered near the threshold; widening the hysteresis band solved it. And a firmware update bricked eleven devices because the update was not atomic and a site lost power mid-rollout; the replacement mechanism writes to a second partition and switches only after verification, and the eleven site visits cost more than the entire storage bill for the year.

What made it work. The local alarm rule at the gateway, which means the system's core function does not depend on connectivity. The device buffer, which turned three network outages into three delays rather than three compliance gaps. And deciding the retention tiers before deployment rather than after, which is the difference between a predictable running cost and an escalating one.

19. Frequently asked questions

Do we need a gateway?

Yes if devices speak a local protocol, if you need decisions faster than a network round trip, if the system must function during an outage, or if aggregating traffic saves meaningful cost. No if devices have direct connectivity, all decisions can be made centrally, and brief outages are tolerable. The offline-operation requirement is the one that most often forces the answer.

How long should we keep raw data?

Long enough to investigate incidents at full resolution — typically thirty to ninety days — with downsampled aggregates kept for as long as the business or regulator requires, and raw archived to cheap object storage if you may ever need it. Keeping everything hot indefinitely is the default that produces the surprising bill in year two.

Cloud platform or build our own?

Use a platform unless you have unusual requirements. Device identity, provisioning, twins, over-the-air updates and fleet management are large amounts of unglamorous work that platforms have already done, and the parts teams underestimate are exactly the parts platforms provide. Build your own processing and applications on top.

How do we handle devices that go offline?

Distinguish expected from unexpected. A battery device sleeping between transmissions is not offline; a device that has missed several expected reports is. Set the missing-data threshold from the reporting interval, route those alerts to whoever maintains hardware rather than to whoever watches the process, and make sure the dashboard shows data age rather than silently displaying stale values.

What is realistic for battery life?

Entirely determined by transmission frequency, and almost always worse than modelled because real conditions produce more transmissions than the model assumed. Measure it in a pilot under realistic conditions before committing to a fleet, and expect threshold-triggered transmissions to be considerably more frequent than the trigger rate suggests.

How do we secure devices in physically accessible locations?

Assume individual devices will be compromised and bound the consequence. Unique credentials per device stored in hardware, secure boot, no debug interfaces enabled, mutual authentication, and platform authorisation scoped so a device can only publish and read its own data. Physical tamper resistance helps at the margin; containment is what actually limits the damage.

When should processing happen at the edge?

When latency requirements rule out a round trip, when the function must survive an outage, when bandwidth costs make sending everything uneconomic, or when data cannot leave the site for regulatory reasons. Otherwise centrally, because central processing is far easier to update, monitor and debug than logic distributed across two hundred gateways.

What is the most underestimated part of an IoT project?

Device management — provisioning, configuration, updates, diagnostics and decommissioning across a fleet nobody can physically reach. It is invisible in a pilot with twenty devices and dominates the work at a thousand. Closely followed by field operations cost, since a single site visit routinely exceeds the value of the device being visited.

Key takeaways

  • Power decides the architecture for battery devices — transmission frequency constrains everything downstream.
  • Buffer on the device so outages become delays rather than permanent gaps.
  • Jitter reporting schedules to avoid correlated bursts, especially after recovery.
  • Tier your retention deliberately. Raw-forever is a decision, and an expensive one.
  • Local decisions must work offline. Anything safety- or operations-critical cannot depend on the network.
  • Device management is the real project. Build it for ten times the fleet you have today.

The data path in an IoT system is the easy part, and it is what every architecture diagram shows. The difficult part is everything around it — getting devices identified, configured, updated, monitored and eventually retired, across hardware in places nobody visits. Teams that design for that from the beginning end up with systems that still work in year three.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch