R&D COPILOT
ROLet’s talk

IoT monitoringHow-to guide

Telemetry that survives outages: buffering, MQTT QoS and duplicate control

A monitoring chart with a gap is worse than it looks. Nobody knows whether the value was fine, the sensor was off or the link was down, and a report built on that chart cannot be defended. We design the telemetry path so that an outage delays data instead of losing it: the gateway keeps readings, the transport confirms delivery, the platform recognises repeats and every reading keeps the time it was measured, not the time it arrived.

By R&D COPILOT5 min read

Name the failure modes before you design for them

Three things interrupt telemetry on a real site. Power drops at the sensor or gateway, sometimes for seconds, sometimes for a weekend. The link drops: a cellular cell is congested, a LoRaWAN gateway loses its backhaul, a Wi-Fi access point is rebooted by IT. And the receiving side restarts: a broker update, a database migration or a cloud region incident. Each one needs a different defence, and a design that only thinks about the link will still lose data on the first power cut.

Write the failure modes in the architecture report with the expected duration of each. A gateway that can hold two hours of data is fine for a flaky link and useless for a long weekend without power at a remote point.

Store and forward on the gateway

The gateway is the right place to protect data, because it sits between sensors that cannot store much and a network that cannot be trusted. Each reading is written to local non-volatile storage first, with a sequence number and the measurement timestamp, and is only removed after the platform has confirmed it. When the link returns, the backlog is sent oldest first, at a controlled rate, so the recovery does not flood the network or the platform.

Size the buffer from the message rate and the longest outage you want to survive, then add margin. Decide what happens when it fills: keep the newest data and thin out older readings to averages, or keep everything and stop accepting new data. That is a business choice and should be written down, not left to a library default.

Choose MQTT QoS per data stream

MQTT defines three delivery levels. QoS 0 sends a message at most once with no acknowledgement, which is fine for high-rate values where the next reading replaces the lost one. QoS 1 delivers at least once: the sender repeats until it receives an acknowledgement, so nothing is lost but duplicates are possible. QoS 2 delivers exactly once through a four-step handshake, at the cost of more traffic and latency on every message.

For most monitoring data we use QoS 1 together with idempotent ingestion, described below. It is simpler and lighter than QoS 2, and the duplicate handling has to exist anyway because the gateway buffer can also resend. Persistent sessions and a broker configured with persistence keep queued messages through a broker restart; Eclipse Mosquitto, for example, documents persistence and limits on queued messages per client.

Make ingestion idempotent

A reading that arrives twice must be stored once. Give every reading an identity that does not change on resend, typically device id plus sequence number or device id plus measurement timestamp, and make the platform treat a repeat as a no-op. Some time-series databases already merge points with the same series and timestamp; InfluxDB, for example, documents that a second point with the same measurement, tag set and timestamp merges with the first. Know which behaviour your store has and do not rely on it accidentally.

The same rule applies further down the chain. If a threshold alert opens a maintenance order in the ERP, a duplicated reading must not open a second order. The integrations layer in RDCopilot uses stable keys for exactly this reason, so a retry is safe at every hop.

Keep time honest

Late data is only useful if it carries the right time. Gateways should synchronise their clocks over NTP when online and, where needed, from GNSS on remote masts. Sensors without a reliable clock can be stamped by the gateway on receipt. Store both the measurement time and the arrival time; the difference shows how late data was and helps explain unusual patterns.

Rules and dashboards need to handle data arriving out of order. An alert rule that evaluated an hour with missing data should be able to re-evaluate when the backlog arrives, and a chart should show the period as delayed rather than silently normal.

Measure completeness, then test recovery

You cannot improve what you do not count. For each device and day, the platform compares readings expected with readings received, and reports completeness, loss, late arrivals, timestamp drift and values flagged as suspect. These indicators appear in the monthly report next to the measurements, so anyone reading the numbers can also see how much to trust them.

Before go-live and after every major change, we run a recovery test on the bench and on site. It is short, repeatable and recorded in the test protocol.

  • Disconnect the gateway uplink for a defined period and confirm the backlog arrives in order
  • Cut gateway power during the outage and confirm buffered data survives the restart
  • Restart the broker and the ingestion service while messages are in flight
  • Resend a batch deliberately and confirm no duplicate readings, alerts or ERP orders appear
  • Compare expected and received counts and record completeness for the test window

What you get when we run it

Reliable telemetry is part of the standard IoT monitoring service, not an extra. Gateways with store and forward, an MQTT broker configured for persistence, idempotent ingestion into a time-series store, completeness indicators in dashboards and reports, and the recovery test protocol are all included. The same data then feeds alerts in the mobile app, maintenance work in the ERP and the monthly reports, with the confidence that a gap means something real.

A one-off implementation fee covers architecture, installation support and configuration. A monthly subscription covers EU hosting, updates, monitoring and support, and extra work is billed at a fixed hourly rate.

Follow the references

Sources & inspiration

Lumos

Devpost project by Dinesh Fatehpuria, Kunal Kislay, Aakash Asim Roy, Sonali Rana

Generator telemetry is queued on the gateway and delivered when the network returns.

This independently created project is credited as inspiration. The workflow and implementation guidance in this article are RDC’s analysis.

Put the guide to work

Start with your workflow.

Tell us what your team needs to do, which systems are involved and where the current process slows down.