All posts

IoT systems · 15 min read

Building an IoT fleet management platform

An end-to-end platform for monitoring and controlling industrial compressors worldwide - from embedded C++ on Yocto Linux all the way to live dashboards, firmware rollout flows, and telemetry pipelines.

At my employer, I worked on a fleet platform that had to do two very different jobs at once. It had to feel immediate - charts moving, commands returning, firmware progress updating live - and it also had to behave like industrial software, where devices disappear for hours, networks are unreliable, and the safe choice is often the slower one.

The interesting part was not any one layer in isolation. It was the seams between them: an internal MQTT bus on the controller, a C++ bridge translating that into Azure IoT Hub semantics, a Python/FastAPI backend and React frontend, worker queues for long-running operations, WebSocket fan-out for live views, Timescale-style telemetry storage for history, and Databricks pipelines for analytical models further downstream.

The shape of the system

At a high level, the platform split into three paths: a control path for commands and firmware rollout, a live path for dashboards, and a data path for storage and analytics.

controller MQTT bus
    -> C++ bridge on Yocto
    -> Azure IoT Hub
        -> Web PubSub groups -> live browser views
        -> worker queues     -> commands / FOTA orchestration
        -> raw storage       -> historical APIs + Databricks bronze/silver/gold

That split was deliberate. I did not want the dashboard request path waiting on analytics jobs, and I definitely did not want a firmware rollout sharing the same assumptions as a chart refresh. The same telemetry can flow to multiple places, but each path should optimise for its own failure mode.

Why the controller bridge mattered

The embedded side ran as a C++ service on a custom Yocto Linux image. Its job was not to be a giant business-logic brain. Its job was to be a narrow, reliable adapter between the controller's internal MQTT ecosystem and the cloud contract.

The bridge subscribed to the controller bus, collected ordinary point updates into grouped telemetry, and published them to IoT Hub on a polling cadence. For time-sensitive cases it had a fast-track path that bypassed batching entirely. That seems small, but it is one of those decisions that decides whether your cloud bill is sane and whether your live charts feel alive.

For the newer controller path, the bridge serialised telemetry into a FlatBuffers DeviceData payload before sending it up. For the older controller path, messages were already JSON and the downstream data platform routed them by message type. In both cases, the device was treated as the source of truth for timestamps; the cloud side deduped and normalised, but it did not pretend it could re-invent device order after the fact.

The same service also auto-detected whether it should authenticate with a shared access key or an X.509 certificate. If a certificate was present, it monitored expiry, generated a CSR with the existing private key, sent that CSR as a tagged telemetry message, accepted the renewed certificate through a direct method response, and atomically swapped the file on disk before reconnecting. That last part matters more than it sounds: a half-written certificate on a field device is the sort of bug you only make once.

// desired properties driving the bridge
{
  "desired": {
    "polling": 30,
    "points": [{ "ids": [1018, 1114], "updateRate": 30, "fastTrack": false }],
    "fastTrack": { "startAt": 1723800000, "duration": 900 },
    "cloudLogs": { "startAt": 1723800000, "duration": 900 },
    "fota": {
      "catalog": {
        "version": "x.y.z",
        "uri": "<signed-download-url>",
        "forceUpdate": false
      }
    }
  }
}

A pattern I like here is that the bridge stayed mostly state-based. It reported what auth mode it was currently using, what points were missing, and what certificate expiry it saw. That makes the cloud side easier to reason about than an event-only design where you have to infer current state from a stream of partial facts.

Twins for intent, direct methods for immediacy

One of the biggest architectural choices was deciding when to use device twins and when to use direct methods. We used both, but for very different jobs.

mechanismbest atfailure modewhere I used it
desired twin propertiesdurable intentarrives late, but survives offline devicessubscriptions, log windows, firmware targets
direct methodsimmediate request/responsefails fast if the device is offlineread/write actions from the UI

Remote read/write actions in the web app went through a worker queue and then into IoT Hub direct methods. The worker built FlatBuffersReadDirectRequest or WriteDirectRequestpayloads, base64-encoded them, invoked the device method, decoded the response, and persisted the result for the UI. That is the right model for something a human just clicked and expects an answer for now.

Firmware rollout was different. The desired firmware version and URL lived in the twin because the device might be asleep, disconnected, or on a poor link. A twin patch is not immediate, but it is durable. That mattered much more than low latency.

A rule that held up

If the operation meant “do this when you can”, I wanted it in the twin. If it meant “do this right now and tell me what happened”, I wanted a direct method. Mixing those two semantics is how you end up with a control plane that is confusing both for users and for operators.

Live dashboards without broadcasting the whole fleet

The frontend was a React app, the backend was FastAPI, and the live path used Azure Web PubSub. The important part was not just that we used WebSockets. It was how narrowly we scoped them.

The browser opened one WebSocket connection using the Web PubSub JSON protocol. From there, it subscribed itself into groups shaped like{hub}_{device}_{stream}: telemetry for one device, logs for one device, FOTA progress for one device, and a broader environment group for cache invalidation events. That let us push precisely what a screen needed instead of turning the whole fleet into a global broadcast problem.

group = `${hubName}_${deviceId}_${stream}`

Telemetry     -> live charts
SystemLog     -> diagnostics tail
FotaProgress  -> rollout progress table
env_<id>      -> invalidate affected queries

The backend exposed a small endpoint that added or removed a connection from one of those groups. The log streaming path was especially neat: IoT Hub routed cloud logs into blob storage, an Azure Function watched the cloud-log container, parsed the JSON lines, grouped them by device, and only pushed them to Web PubSub if that device group actually had listeners.

That meant the expensive path was demand-driven. If nobody was watching a device's logs, the system did almost nothing beyond storing them. If someone opened the diagnostics view, the same raw feed became a live tail.

The five-second race I kept

Web PubSub gives the browser a connection ID before the rest of the system is fully caught up. In practice I had a real race between “frontend wants to join groups now” and “backend has finished putting this user into the right environment context”. The fix in both backend and frontend was an explicit five-second delay. It is not elegant, but it is honest: sometimes distributed systems hand you eventual consistency, and pretending otherwise just moves the bug somewhere harder to debug.

Firmware updates on bad networks

Firmware-over-the-air is where the difference between SaaS software and industrial software becomes painfully obvious. On a laptop, the answer to a failed update is often “download it again”. On a compressor in the field, the answer has to be “make sure the state machine still makes sense tomorrow”.

The rollout flow started in the backend, but it did not execute there directly. The backend persisted intent, approval state, and device targets, then pushed work onto Service Bus. A worker app published the chosen firmware into a distribution container, generated signed download URLs, and updated device twins instead of trying to keep a long-lived request open.

Two choices here turned out to matter a lot. First, the actual device state came back from twin sync, not from wishful thinking. The worker periodically compared desired firmware, reported firmware, OTA state, and the currently assigned download URL. That is how it decided whether a device was still pending, already in progress, complete, superseded by a newer version, or just failed.

Second, the signed URLs were treated as short-lived infrastructure, not permanent identity. The worker published firmware with a limited SAS lifetime, renewed URLs before expiry, and then patched twins again in batches. The batch size was capped, and the renewal jobs were createdsequentially because IoT Hub twin update jobs do not behave well if you flood them in parallel.

for batch in chunks(devices, 100):
    job = create_twin_update_job(batch, {
        firmware_url: refreshed_signed_url,
        firmware_version: target_version
    })
    wait_until_job_finishes(job)   # create next batch only after this one ends

That is a good example of an industrial tradeoff. The “fast” design was to fan out every job immediately. The correct design was to respect the platform's throttling behaviour and move more slowly. The retry layer in the worker even handled HTTP 429 with exponential backoff and jitter because quota management is not theoretical in a real fleet.

For the MK6 controller path, there was a second firmware mechanism: catalog updates written into desired twin properties, which the controller-side bridge translated back onto the internal bus. That let the cloud express intent in one place while the embedded side kept local ownership of how a safe update is actually applied.

Telemetry does not stop at the dashboard

One thing I liked about this platform is that it did not force one storage system to do every job badly. The hot path and the analytics path were different on purpose.

For the web app's historical APIs, telemetry landed in separate time-series databases by device family. The schema was deliberately unusual: a message table for envelope metadata, then point-specific hypertables like point_XXXX_telemetry. That only works when your point catalogue is known, but for a fixed industrial domain it makes “give me point 11402 for this device over this range” very cheap, and that is exactly the query shape the dashboards used.

The deeper analytics path ran through Databricks. There the pipeline was explicitly medallion-shaped, but with one practical twist: ingest was a two-step process because the first stage ran on classic compute with storage credentials and the second stage ran on serverless compute reading from a Unity Catalog external location.

IoT Hub raw files
  -> staging Delta path
  -> bronze.messages
  -> silver typed tables
  -> gold marts for connectivity, firmware, and recent telemetry

The NanoController pipeline decoded the base64 IoT Hub body, extracted message-level fields, and stored the raw JSON in bronze. Silver then reparsed only the relevant time window and routed rows into five typed tables: commission, events, heartbeat, system, and telemetry. It also handled schema drift in sensor payloads and normalised legacy device ID variants so later joins would not quietly fragment the same physical machine into two logical ones.

The MK6 path was different because the payload was FlatBuffers rather than JSON. Bronze kept the raw message body and content type. Silver installed the shared schema package, deserialised only the telemetry content type, exploded measurements into point rows, and deduped on device, timestamp, and point number. Again, the theme was state over stream mythology: I never assumed exactly-once delivery, so dedupe was a first-class part of the design.

On top of that, dbt built gold tables for questions the product and reporting layers actually cared about: recent telemetry joined with device metadata, per-device connectivity profiles, monthly connection statistics, and latest firmware versions seen in system messages. That gave us a clean boundary: the operational product path stayed fast, and the analytical path stayed expressive.

One serverless constraint that changed the design

In one of the telemetry transforms, normal Spark caching was not available on the serverless runtime we were using. The practical workaround was to materialise a temporary Delta table and let five downstream transforms reuse that instead of each re-parsing the raw JSON. It is not the design I would have invented on a whiteboard, but it was the right design for the platform I actually had.

What I would change

The sharpest edge is that the platform still has a lot of modelled knowledge about fixed point numbers and device-specific schemas. That is great for performance and terrible for change velocity. It works because industrial telemetry is relatively stable, but every new device family reminds you how much plumbing is implicit in that decision.

I would also like a cleaner unification between the operational time-series storage and the analytical lakehouse path. Right now that split is justified - the query patterns are genuinely different - but it does mean two places to reason about retention, dedupe, and schema evolution. The architecture is correct; it is just not cheap in cognitive load.

Finally, I would replace the explicit timing gaps around real-time connection setup with stronger end-to-end acknowledgement semantics. The current version works, and sometimes that is enough, but it still bothers me whenever I see a deliberate sleep in a production control path.

The stack

I like this project because it forced me to think across the whole stack. It was not enough for any single layer to be elegant. The system only worked because the device, cloud, UI, and data platform made compatible promises to each other.