ESP32 Provisioning and MQTT Reconnects: Offline Is Normal

ESP32 firmware has to treat offline as normal: Wi-Fi provisioning that does not fight the radio, capped reconnects with backoff, an MQTT outbox that expires in 30 seconds, and a local gateway.
TL;DR: ESP32 Wi-Fi provisioning and MQTT reconnect logic are where IoT firmware actually breaks, so design them before the happy path. In 2026 that means ESP-IDF's network_provisioning component over Bluetooth LE with Security 2, a capped Wi-Fi retry count that falls back to provisioning, a jittered backoff the libraries do not give you, a re-subscribe on every MQTT connect, an outbox you size on purpose because it forgets after 30 seconds, and a local gateway broker that keeps working when the internet does not.
Why Should ESP32 Firmware Treat Offline as the Normal State?
ESP32 firmware should treat offline as normal because every link in the path drops: the access point restarts, the broker restarts, the uplink goes down. ESP-IDF says it plainly: "It is the application's responsibility to reconnect." Firmware written for the connected case spends its life in the disconnected one.
I learned this building an end-to-end IoT smart-farming solution: C/C++ firmware on ESP32 devices, a Raspberry Pi as the local gateway, a Django backend and a Vue.js dashboard. It was a proof of concept for indoor farming, so the devices used Wi-Fi and ran on integrated batteries. About five devices posted sensor readings every 10 seconds, each with a unique device ID, mostly through a custom API on the Django backend, with MQTT running through a custom broker alongside it. Readings also showed on an OLED module on each device. We experimented with LoRa for wider area coverage and with ESP-NOW between devices.
The devices never went down. The part that failed most easily was getting them onto the network in the first place.
Why Is Wi-Fi Provisioning the Hardest Part of ESP32 Firmware?
Wi-Fi provisioning is the hardest part because the device has to be an access point and a client at the same time on one radio. To take credentials, it hosts a hotspot; to test them, it scans and connects to the farm's network; and it has to report the result back to a phone that is connected to that hotspot.
On my build, that setup routine was the fragile path. The device hosted the hotspot, scanned for networks, tried the connection and reported back, and the only way to keep it responsive was multi-core programming: I offloaded the setup task to the ESP32's second core. When a device lost Wi-Fi later, it went back into AP mode and reconnected automatically when the network returned, with reboot logic as the last resort.
ESP-IDF's own provisioning documentation describes exactly that tension. With SoftAP provisioning, "The device uses the same radio to host the SoftAP and also to connect to the configured AP. Since these could potentially be on different channels, it may cause connection status updates not to be reliably received by the phone." The phone also "has to disconnect from its current AP in order to connect to the SoftAP." Hand-rolled setup code inherits every one of those problems.
How Should an ESP32 Be Provisioned in 2026?
Provision over Bluetooth LE with ESP-IDF's network_provisioning component and Security 2. As of ESP-IDF v6.0, "The ESP-IDF component wifi_provisioning has been removed from ESP-IDF and is supported as a separate component. It has been renamed to network_provisioning", so new firmware pulls it from the component registry.
Bluetooth LE takes the radio conflict away. ESP-IDF's docs: it "has the advantage of maintaining an intact communication channel between the device and the client during the provisioning, which ensures reliable provisioning feedback," and the phone app can find and connect to the device without leaving the app. The cost is memory: "about 110 KB memory at runtime", and "almost all the memory can be reclaimed" if Bluetooth is not needed after provisioning. SoftAP stays an option for devices without Bluetooth LE.
# main/idf_component.yml
dependencies:
espressif/network_provisioning: "^1.3.1"network_prov_mgr_config_t config = {
.network_prov_wifi_conn_cfg = {
.wifi_conn_attempts = 5, // then erase credentials and restart provisioning
},
.scheme = network_prov_scheme_ble,
.scheme_event_handler = NETWORK_PROV_SCHEME_BLE_EVENT_HANDLER_FREE_BTDM, // reclaim BLE memory afterwards
};
ESP_ERROR_CHECK(network_prov_mgr_init(config));
bool provisioned = false;
ESP_ERROR_CHECK(network_prov_mgr_is_wifi_provisioned(&provisioned));
if (!provisioned) {
ESP_ERROR_CHECK(network_prov_mgr_start_provisioning(
NETWORK_PROV_SECURITY_2, (const void *) &sec2_params, service_name, NULL));
}wifi_conn_attempts is the managed version of my AP-mode fallback: the official example's default is 5 attempts, after which "Provisioned credentials are erased and internal state machine is reset", so a mistyped password sends the device back to setup instead of retrying forever. A value of 0 keeps the "legacy behavior of infinite connection attempts."
Security 2 is "SRP6a-based shared key derivation and AES256-GCM mode encryption of the data", and in production the username and password are not embedded in firmware but given to the user "by suitable means, e.g., QR code sticker." ESP-IDF v6.0 also turned Security 0 and 1 off by default, so the secure option is now the default one.
How Should ESP32 Firmware Reconnect Wi-Fi and MQTT?
Reconnect Wi-Fi yourself with a capped, jittered backoff, and let MQTT reconnect on top of it, re-subscribing on every connect. Neither library backs off for you: ESP-IDF's Wi-Fi guide offers a retry counter or "reconnect immediately in the first N continuous reconnection, then give a delay", and the ESP-MQTT client retries on a fixed reconnect_timeout_ms, 10 seconds by default.
A fixed interval is fine for one device and bad for a fleet: when an access point restarts, every device behind it retries in lock step. A sketch of the Wi-Fi side:
static void on_network_event(void *arg, esp_event_base_t base, int32_t id, void *data)
{
if (base == WIFI_EVENT && id == WIFI_EVENT_STA_DISCONNECTED) {
if (leaving_on_purpose) return; // esp_wifi_disconnect() raises this event too
if (attempts++ < FAST_RETRIES) {
esp_wifi_connect(); // first N: reconnect immediately
} else {
uint32_t backoff_ms = MIN(60000, 1000u << MIN(attempts - FAST_RETRIES, 6));
esp_timer_start_once(reconnect_timer, (backoff_ms + esp_random() % 1000) * 1000ULL);
}
} else if (base == IP_EVENT && id == IP_EVENT_STA_GOT_IP) {
attempts = 0;
}
}The leaving_on_purpose check comes from ESP-IDF: if the disconnect came from esp_wifi_disconnect(), "the application should not call esp_wifi_connect() to reconnect." And every socket dies with the link: "the default behavior of LwIP is to abort all TCP socket connections on receiving the disconnect."
On the MQTT side, two defaults need attention:
- Subscriptions do not survive a reconnect. Clean session is on by default, and MQTT 3.1.1 says such a client "has to subscribe afresh to any topics that it is interested in each time it connects." Espressif's own example subscribes inside
MQTT_EVENT_CONNECTED, which runs on every connect. Do the same. keepalive = 0does not disable keep-alive. ESP-MQTT's header: setting it to 0 "doesn't disable keepalive feature, but uses a default keepalive period" of 120 seconds. Set the value you mean.
What Happens to ESP32 Telemetry While MQTT Is Offline?
While the ESP32 is offline, QoS 0 telemetry is lost and QoS 1 telemetry waits in a RAM outbox for 30 seconds, then disappears. ESP-MQTT's docs: "QoS0 publish fails when disconnected; QoS1/2 are stored in the outbox until ACK", and "After CONFIG_MQTT_OUTBOX_EXPIRED_TIMEOUT_MS messages will expire and be deleted."
That makes QoS 1 a retry mechanism, not an offline buffer. Three details from the ESP-MQTT v1.1.0 source and docs:
- Deletion is silent by default. The
MQTT_EVENT_DELETEDevent only fires ifCONFIG_MQTT_REPORT_DELETED_MESSAGESis enabled, and its default is off. - The outbox has no size limit by default.
outbox.limitis a byte budget, and the library only checks it when it is above 0. On an unstable link Espressif warns that messages pile up and "may pose a problem if the outbox size becomes too big over time." - The outbox lives in RAM. A reboot or
esp_mqtt_client_stopempties it. For anything that must survive, Espressif points to a custom outbox: "a specific implementation of message outbox is needed (e.g. persistent outbox in NVM or similar)."
const esp_mqtt_client_config_t mqtt_cfg = {
.broker.address.uri = "mqtt://gateway.local",
.credentials.client_id = device_id, // provisioned, not derived from the MAC
.session.keepalive = 30,
.outbox.limit = 32 * 1024, // bytes; 0 means no limit
};My proof of concept sent readings straight out and showed them on the device; we experimented with saving data on flash. For production telemetry I'd timestamp every reading on the device, write it to flash before sending, and delete it only after the broker acknowledges it, so a long outage costs latency, not data. The client ID matters for the same reason: the default is ESP32_%CHIPID%, built from "the last 3 bytes of MAC address", so give each device the unique ID it was provisioned with.
Where Does a Gateway Fit Between ESP32 Devices and the Backend?
A gateway gives the devices a broker on the local network, so a broken internet uplink stops being their problem. On my build the Raspberry Pi was that gateway. In 2026 I'd run Mosquitto on it and bridge to the cloud broker, letting the gateway, which has disk and memory, absorb outages the ESP32 cannot.
Mosquitto's bridge already does what the ESP32 libraries leave to you. restart_timeout takes a base and a cap and uses "a backoff mechanism based on "Decorrelated Jitter"", and a bridge's cleansession defaults to false, which "means that all subscriptions on the remote broker are kept in case of the network connection dropping." Two defaults to change: persistence "Defaults to false", so queued data lives in memory only, and the docs now recommend "plugin based persistence" over the built-in file.
# /etc/mosquitto/conf.d/bridge.conf on the gateway
persistence true
persistence_location /var/lib/mosquitto/
connection cloud
address broker.example.com:8883
bridge_cafile /etc/ssl/certs/ca-certificates.crt
topic farm/# out 1
cleansession false
restart_timeout 5 300Devices out of Wi-Fi range are where the LoRa and ESP-NOW experiments fit: a different radio link to a node that does have a path to the gateway, with the same rule at every hop. Each link assumes the next one is down.
Reconnect logic is the same problem one layer up from the browser, where a MIDI device unplugged mid-note has to be handled as a state, not an error.
Limitations
- The smart-farming build was a proof of concept under NDA: about five devices on Wi-Fi in an indoor setting. No client, location or production data is described here.
- The 2026 recommendations come from ESP-IDF v6.1,
network_provisioning1.3.1, ESP-MQTT v1.1.0, the MQTT 3.1.1 specification and Mosquitto 2.1.2, read on 2026-10-08. They are not what the original firmware ran. - The Wi-Fi backoff and MQTT configuration snippets are sketches built from those APIs and were not compiled for this article. The provisioning snippet follows Espressif's official
wifi_provexample. - Bluetooth LE provisioning needs a phone app; Espressif publishes Android and iOS apps and a command-line tool,
esp_prov.
Design the disconnected path first and the connected one comes almost free. What does your firmware do with a reading it cannot send? More field notes on embedded and IoT.
References
- ESP-IDF v6.0 migration guide: Provisioning and Unified Provisioning: transports, Security 2, memory.
- network_provisioning component and its wifi_prov example.
- ESP-IDF Wi-Fi driver: disconnect events and reconnection.
- ESP-MQTT (v1.1.0): reconnect, outbox, expiry and configuration.
- MQTT 3.1.1 specification: clean session and subscriptions.
- Mosquitto
mosquitto.conf(2.1.2): bridges,restart_timeout, persistence.