OTA firmware updates provide signed firmware images to devices over a network connection rather than manual reflashing via physical access or cables. When performed correctly, the system rests on three pillars that should never be sacrificed: cryptographic signing of every package, atomic or A/B installation so no device is ever left half flashed, and staged rollout with live telemetry so a bad release can be stopped before it reaches the entire fleet. Neglect one, and you are one broken build away from destroying your deployment.
TL;DR:
- Staged rollouts combined with telemetry allow teams to catch bad updates early and dramatically reduce bricking risk.
- Secure signing key management, including hardware protection, rotation, and audit logging, helps prevent unauthorised firmware from ever running.
- Atomic installation techniques with layered rollback prevention ensure devices update to new content or nothing at all.
- Metadata must include product IDs, versioning, hashes, and dependency rulesso incompatible or older firmware is never flashed.
- Secure transport such as TLS with certificate pinning helps prevent spoofed updates and protects firmware delivery over existing telemetry channels.
Building Robust Infrastructure Software
PODTECH engineers tailor software, automation, and data platforms that help infrastructure stay up for the most demanding enterprise applications.
Learn MoreTable of Contents
- How OTA firmware updates work end to end
- Security best practices for signing and transport
- Choosing an update architecture: A/B, atomic install and rollback
- What belongs in an OTA update package and its metadata
- Rolling out updates safely without risking the fleet
- Device-side design: agents, power-fail safety and recovery
- Standards worth knowing: Uptane, OMA-DM and ISO guidance
- Transport and orchestration patterns that work in practice
- Why enterprise fleets need more than a good bootloader
- The gap between a working demo and a fleet that survives production
- Getting OTA architecture right for critical fleets
- Sources
- FAQ
How OTA firmware updates work end to end
Conceptually, an OTA system consists of five steps: the campaign server, signing authority, metadata service, on-device agent, and the bootloader. The campaign server determines which devices receive which build. Signing, done separately on hardware you control, cryptographically signs the image before it ever reaches a device. Metadata about the image, including version, target hardware, and hash, is described in a package. The device agent downloads or receives that package. Finally, the bootloader or installer applies it.
The flow itself is fairly standardised: announce, download, verify, install, test, then activate or rollback. The device is informed that an update is available either through polling a server or receiving a push message. The update package is downloaded and verified with the signature and hash matching what is trusted by the bootloader. The package is then installed to an inactive partition or memory region and booted into for a preconfigured time period to run tests. If the tests pass or the timer expires, it commits the changes. If not, it automatically rolls back.
Package-based updates and image-based updates solve the same problem with different trade-offs. The package or delta approach only sends the changed bytes necessary to patch from one firmware version to another. That makes sense for resource-constrained devices with limited flash memory, but it increases CPU cost because the device must reassemble the full image. The full image approach is easier to verify and reason about, but it requires enough spare flash to store a second complete copy. That is why A/B partitioning exists: it solves the issue at the platform level instead of treating it as an optional bonus feature. Misjudging this trade-off causes many OTA debugging sessions that end with “no space left” on the device.
Security best practices for signing and transport
All trustworthy OTA designs begin with one principle: never execute unsigned code. Generate a signing key pair. Keep the private key secured in a hardware security module or another secure enclave. Verify the signature on the device before executing a single byte of new firmware. The Secure-by-Design handbook treats OTA signing, anti-rollback, and recovery testing as baseline requirements for every connected product, not special provisions reserved for regulated industries.
Transport security matters just as much as the payload itself. TLS with certificate pinning helps prevent a compromised or spoofed update server from pushing malicious firmware. Streaming the update over an existing authenticated channel, such as MQTT, can eliminate the need to open a second and potentially less scrutinised connection for updates. Texas Instruments recommends exactly this pattern: stream OTA updates over the device’s existing telemetry channel, authenticate each chunk, and use bundle protection so a corrupted stream never reaches the bootloader.
Version binding closes the downgrade loophole that catches many teams. Firmware metadata should include a monotonic version counter that is compared to a value stored in fused or write-once memory on the device. That ensures an attacker cannot replay an older but still signed image to reopen a previously fixed vulnerability.
A few operational controls tie the technical and organisational sides together:
- Signing keys live in an HSM, never in source control or build scripts.
- Release authorisation requires named sign-off before a build enters the signing pipeline.
- Every signing event and rollout decision writes to an immutable audit log.
- Keys rotate on a fixed schedule and immediately after any suspected compromise.
Extra Credit Tip: Manage your signing key rotation schedule the same way you manage certificate expiry: put it on a calendar and assign an owner. A lapsed or leaked key silently breaks your entire chain of trust.
Recent research into connected-vehicle OTA security points toward integrated authentication schemes that reduce cryptographic overhead for constrained electronic control units without sacrificing integrity guarantees. If your devices operate on a tight compute budget, this is worth watching closely.
Choosing an update architecture: A/B, atomic install and rollback
Three architectural patterns apply to most fleets. The right choice depends on your flash budget and your risk tolerance.
- Single-bank with recovery partition. The lowest-cost option in terms of flash: one active firmware area and a separate but small recovery image that can be independently verified. This can work for lower-risk devices, but it creates a real vulnerability window if power is lost during a write.
- Legacy A/B (dual full-image). Two complete firmware slots, one active and one inactive. The update writes to the inactive slot, boots into it for validation, then sets a flag making it active. This roughly doubles flash requirements, but rollback is simple: reset the flag.
- Virtual A/B. Similar to full A/B but more flash-efficient, common on modern embedded Linux and mobile platforms. It shares static data between slots and only copies the dynamic partition. You get most of the benefits of full A/B without fully doubling storage requirements.
Atomic install is the property that makes any of these patterns actually safe: the device either completes the whole install and cleanly switches over, or it never moves from where it started. The mechanism that enforces this is the test-boot window, an automatic timer initiated after flashing a new image. If the device signals that it booted successfully and passed a basic health check within this window, the update is committed. Otherwise, the bootloader automatically rolls back to the previous slot with no user action required.
Rollback controls should be layered. There is no single silver-bullet check that works on its own. A monotonic version counter prevents replay of older images. Server-side eligibility rules prevent the device from even receiving an offer for a build it should not run. Fused write-once counters in silicon provide a hardware-enforced floor that survives even a compromised bootloader. Every layer costs engineering effort. Every layer also blocks a distinct avenue of attack that the others do not.
What belongs in an OTA update package and its metadata
A firmware package is only as secure as the metadata wrapped around it. At minimum, each package should include:
- Product and hardware identifiers so a build for one board revision never lands on another by accident.
- Semantic version and monotonic version counter, including a human-readable version and an anti-rollback check resistant to tampering.
- Cryptographic hash of the payload, verified separately from signature verification to detect corruption during transmission.
- Who signed it, linking the package back to a key and therefore indirectly to a specific release approval.
- Dependency rules, such as the minimum bootloader or prior firmware version that must already be present before installation.
Eligibility enforcement happens by comparing this metadata to the device’s own state before the download begins. It is not an after-the-fact check. A device that identifies itself as hardware revision B should never be presented with an update package built and signed for revision A, and metadata is where that rule is enforced.
Chunked and resumable downloads become just as important as the security fields on low-flash devices. If partially stored images are handled securely, including per-chunk hash checkpoints, then a dropped Wi-Fi connection or a power blip during download does not force a complete restart. That removes one of the biggest causes of stalled fleets in the field.
Rolling out updates safely without risking the fleet
Staged rollout is the operational equivalent of A/B testing and should be handled with the same discipline: pilot, canary, ramp, then full. A pilot group of a few internally controlled devices catches obvious issues. A canary group made up of real production devices but kept intentionally small surfaces problems that only appear at scale or under field conditions. A ramp phase increases exposure gradually, often doubling the percentage of the fleet at each interval. Full rollout should happen only after each earlier gate remains error-free for a predefined period.
Telemetry determines whether each gate opens or closes. Watch for:
- Crash-free boot rate immediately after update, compared with the previous firmware baseline.
- Rollback rate, meaning how often the test-boot window triggers an automatic revert.
- Connectivity and check-in success rate in the hours following installation.
- Battery or power-draw anomalies on the new build versus the old one.
Tip: Make your halt trigger deterministic rather than subjective. If rollback rate exceeds a specified threshold during the canary window, automatically place the campaign on hold. No drama required.
Every release should have an audit trail from the signed package back to both the vulnerability ticket or feature request that motivated it and the named approval that authorised signing. Texas Instruments shows how this can work in practice through job orchestration systems that report per-device status back through the same pipeline that issued the update, which greatly simplifies audit reconstruction after an incident.
Device-side design: agents, power-fail safety and recovery
Responsibilities of the on-device OTA agent go far beyond “download and flash.” The agent must schedule installs around the device’s actual workload, skip installation during safety-critical operations, independently verify signatures regardless of any server-side checks, and store the downloaded image in a location trusted by the bootloader before initiating installation.
- Trust no one downstream. The agent verifies the signature and the hash itself. It does not assume some earlier hop in the network already validated the payload.
- Delay smartly. If an agent installs mid-cycle on a safety or production asset, that is poor design, not bad luck. Define clear no-interruption windows.
- Respect the watchdog. Test-boot windows must cooperate with hardware watchdog timers. If firmware freezes during validation boot, the watchdog should kick the device back rather than let it hang forever.
- Fall back to a signed recovery image you trust. If both slots fail to boot, the device should still be able to load minimal firmware signed separately from the primary image and reconnect to the network for repair.
Run a predetermined test matrix instead of an ad hoc smoke test before fleet-wide rollout. Abort installation at different stages, halfway through download, during write, immediately after write, and during boot test, then confirm the device returns cleanly each time. Attempt explicit downgrades to verify rollback protection actually works. Run every failure case on multiple real devices, because flash wear and manufacturing variation can change power-fail behaviour in ways a single test unit will never reveal.
Standards worth knowing: Uptane, OMA-DM and ISO guidance
Uptane’s approach to separate director and image repositories can be useful well beyond automotive systems. The director repository tells each ECU what to run, while the image repository stores the cryptographically verified binaries. Offline signing keys can protect the image repository even if an attacker compromises the internet-connected director system. This is especially valuable when one campaign must deliver updates to many different compute targets within a single fleet.
OMA-DM is broader and includes device configuration and diagnostics as well as firmware delivery. It is useful when OTA is only one part of a larger device management problem. If you are working on safety-critical or closely automotive-adjacent projects, ISO/SAE 21434 defines process expectations around cybersecurity engineering, while ISO 24089 applies specifically to software update engineering.
Heavyweight standards make sense when you already have real multi-party trust boundaries, such as separate teams signing separate artefacts, or when regulatory exposure is high. A single-product IoT device with one signing authority usually does not need the full weight of Uptane. It does need the same principles, sized appropriately.
Transport and orchestration patterns that work in practice
Reusing a telemetry connection to deliver OTA updates avoids a second authentication handshake, but update packets now compete with telemetry traffic for bandwidth. That makes chunk sizing and rate limiting more important than they would be on a dedicated channel. A separate HTTPS download path is often easier to reason about and firewall independently.
Vendor-specific terminology aside, cloud orchestration looks conceptually similar across providers: jobs are your campaign, device groups are your target population, and staggered job execution is how you enforce staged rollout discipline. Particle’s cloud documentation shows atomic updates, encrypted transport, and staged releases as a practical example teams can learn from and adapt rather than copy blindly. Whatever tooling stack you choose, the bootloader hook that vetoes applying a new image unless it is known to be good is the control that links transport decisions back to device-level safety.
Why enterprise fleets need more than a good bootloader
Correct cryptography and rollback logic are required, but they are rarely enough on their own at fleet scale. The difference between a functioning prototype and a production-grade OTA system is the combination of technical controls, operational discipline, and a release process that can survive years of real traffic.
The gap between a working demo and a fleet that survives production
Rollback is an afterthought on too many OTA projects. Build it first and test it last. Before you ever write the happy-path installer, test interrupted installs, downgrade attempts, and watchdog behaviour. Recovery from bad installs is what determines whether your fleet survives its first bad release. Unsigned packages and skipped staged rollouts are the two failures that show up in almost every post-incident report.
— Harry
Getting OTA architecture right for critical fleets
The happy path rarely needs much work. Most teams get that part right. The difficult part is the part you never see in the demo: rollback under real power-fail conditions, staged rollout gates that actually stop bad releases from progressing, and audit trails that stand up to compliance teams when questions arise. PODTECH builds that layer through Master Systems Integration for fleets that already have BMS, PMS, or NMS platforms and need OTA wired into existing telemetry, and through Product Development engagements for teams building updateable devices from scratch.
Work under Legacy Modernisation and SaaS Development covers these bases respectively. If your fleet is large-scale, safety-critical, or burdened by compliance requirements that a bench-tested bootloader will not satisfy, get in touch via PODTECH and we can scope the engagement.
Sources
A More Secure and Reliable OTA Update Architecture for IoT Devices — Texas Instruments
Available Under CC BY 4.0 license: https://www.ti.com/lit/wp/sway021/sway021.pdf?ts=1617119077461&ref_url=https%3A%2F%2Flearning.oreilly.com%2F
PDF | 8 Pages | July 2021 | TI
As updates to IoT devices are projected to become increasingly frequent and complex, Texas Instruments recognised a need to reimagine how these updates can be pushed out securely and reliably. To that end, they designed a more robust OTA update architecture by splitting the payload into four chunks and encrypting them with secure envelopes.
This paper walks through a detailed example of pushing an OTA update to a TI CC3220ST Wi-Fi chip. Along the way it covers key concepts like manifests, bootloader stages, security considerations, rollback features, and more.
Supported Devices:
TI CC3220ST Wi-Fi chip
FAQ
Are OTA updates necessary?
Yes. Any connected device that will need security updates or feature patches after shipping needs OTA capability. Without OTA, the only way to patch a known vulnerability is a physical recall or a service engineer visit, which is almost never feasible once your fleet grows beyond a few hundred units.
What devices use OTA firmware updates?
Nearly any embedded device connected to a network can use OTA firmware updates: routers, industrial monitors and sensors, building telemetry devices, vehicles, medical equipment, and consumer IoT products. The common requirement is that the device becomes difficult or expensive to access directly once deployed.
Is the OTA update process safe?
It is safe if it is built with signed packages, atomic or A/B installation, and staged rollout with monitoring. Those are the same three controls discussed throughout this guide. It becomes unsafe the moment you skip one, especially unsigned packages or a rollout with no canary phase, both of which the Secure-by-Design handbook treats as minimum requirements.
Where can I find the OTA update file, and who manages it?
An update file should never live on a public or unauthenticated endpoint. It belongs on the campaign server or cloud job service that distributes it. Ownership usually sits with the team responsible for the release pipeline. Many organisations do not have that internal capacity and instead engage a partner such as PODTECH to architect and operate it as part of a Master Systems Integration programme.
