A practical guide to embedded software testing — host-based unit tests, on-target testing, hardware-in-the-loop rigs, production test fixtures, and CI pipelines. Lessons from thermometer guns, PLCs, gas detectors, and medical monitors we shipped.
Desktop developers take testing for granted: write a test, run it in milliseconds, watch it go green. Embedded developers live in a harder world. The code runs on a microcontroller with 64 KB of RAM, talks to sensors over I2C, and drives a motor whose stall current can brown out the board. You cannot just run the firmware on your laptop — half of it is the hardware. A bug that is a minor annoyance on a server can be a bricked device, a failed certification, or a factory line stopped at midnight.
Over years of shipping firmware — infrared thermometer guns, STM32-based PLC controllers, ESP8266 gas detectors, and multi-parameter medical monitors — we converged on a layered testing strategy. No single technique covers everything; each layer catches the bugs the others miss. This guide walks through all five layers, with concrete examples from hardware we actually built.
The cheapest bug is the one you catch without touching hardware. A surprising amount of firmware logic has nothing to do with peripherals: temperature compensation algorithms, CRC calculations, protocol parsers, state machines, calibration curves. All of it can run on your development machine.
Take the infrared thermometer gun we developed as a complete production solution (LandwinGUN). The project shipped with a dedicated temperature gun algorithm document — the math that converts raw infrared sensor readings into a body temperature display, compensating for ambient temperature and emissivity. That algorithm is pure computation: no GPIO, no timers, no hardware. We extracted it into a hardware-independent module and tested it on the host with hundreds of input vectors, including edge cases like sub-zero ambients and saturated sensor readings. A framework like Unity or Ceedling makes this natural in C; the discipline is architectural — keep the algorithm separate from the driver that feeds it data.
The rule of thumb: if a function does not touch a register, it should be unit-testable on the host. Teams that follow this rule typically discover that 40 to 60 percent of their firmware logic qualifies, and that is 40 to 60 percent of bugs caught in seconds instead of on the bench.
Host tests prove the logic; on-target tests prove the logic on the actual silicon. Compilers for Cortex-M differ from x86 compilers in the details that matter — integer promotion, struct packing, floating-point behavior without an FPU. A test that passes on your laptop can fail on the chip, and only running on the chip finds out.
Our STM32F407-based PLC industrial control project (STM32 PLC) runs FreeRTOS with a FatFS file system, SDIO storage, SRAM expansion, and multiple UART channels. For a system that complex, we built a small on-target test harness: a dedicated FreeRTOS task that runs at boot (in test builds), exercises each driver — write a file, read it back, verify the CRC — and reports results over UART. The harness stays in the codebase permanently, guarded by a compile flag. When a new engineer changes the SDIO driver, the test task tells them within seconds whether they broke storage.
On-target tests are slower to write than host tests, so spend them where the hardware is involved: driver initialization sequences, interrupt handling, DMA transfers, and timing-critical code. That is exactly the code host tests cannot cover.
Unit tests check organs; hardware-in-the-loop (HIL) checks the whole animal. In HIL, the real firmware runs on the real board, but the physical world is simulated: a test rig feeds scripted sensor inputs and verifies actuator outputs automatically.
The ESP8266-based smart gas detection card (PC_CARD) is a good example of why this matters. The board monitors three gas sensor channels, drives a valve motor with open/close/stall protection, and reports over WiFi. Testing the valve logic by hand means holding a gas source to a sensor and watching a motor — slow, unrepeatable, and smelly. Our HIL rig replaced the sensors with DAC-driven voltages following scripted gas profiles and monitored the motor driver outputs with the rig's own ADCs. A full open-alarm-close-stall cycle ran in under a minute, unattended, and caught a stall-detection timing bug that manual testing had missed for weeks.
HIL rigs cost real engineering time — budget a week or two per product. They pay for themselves the first time a regression would otherwise have shipped. The trick is to design for testability from the schematic: expose test points for every signal the rig needs to drive or observe, and put the board into test mode with a simple, documented command.
Testing does not end when development ends. Every unit that leaves the factory needs a production test: does the PCB work, are all components soldered, is the firmware flashed and calibrated?
The thermometer gun program taught us this twice. The dual-version hardware design (ThermoGun) defined four working modes — default, memory, setting, and calibration. Those modes doubled as production test hooks: the calibration mode, intended for end users, let the factory verify sensor accuracy against a blackbody reference. And the production solution package (LandwinGUN) shipped with complete Gerber, silkscreen, pick-and-place, and stencil files — the manufacturing data that a test fixture designer needs to build bed-of-nails contact with every net on the board.
A production test typically runs in three stages: a boundary/ICT check that every net is connected, a functional test that exercises sensors and actuators (reusing the HIL scripts), and a calibration step that writes per-unit correction values to flash. Skip any of the three and the factory will ship units that your support team gets to debug — at ten times the cost.
Continuous integration works for embedded, with adaptations. Our CI pipeline for firmware projects does four things on every commit:
What CI cannot do (yet) is run the on-target and HIL tests per commit — those need physical boards. The pragmatic setup keeps one or two boards attached to a CI runner for nightly runs: every night, the latest firmware is flashed and the full on-target and HIL suites execute. Developers get host-test feedback in minutes and hardware feedback by morning.
For safety-adjacent products, testing is also a compliance story. The multi-parameter medical monitor we built (Medical Monitor) integrates ECG, NIBP, SpO2, temperature, and respiration monitoring with medical-grade accuracy of ±1%, meeting the YY 97064-2015 standard. For a device like that, the test plan is part of the product: every requirement traces to a test, every test result is recorded, and the record is what the certification body audits. Even for non-medical products, borrowing this discipline — a written test plan, requirement-to-test traceability, recorded results — is the difference between "we think it works" and "we can prove it works."
Our pre-ship checklist, refined over dozens of products:
Embedded software testing is not one technique but five layers, each catching what the others cannot: host unit tests for logic, on-target tests for the silicon, HIL for the integrated system, production tests for manufacturing, and CI to hold it all together across every commit. Teams that invest in all five ship firmware that behaves — on the bench, on the factory floor, and in the customer's hands.
If you are building a product with firmware in it — a sensor device, an industrial controller, a consumer gadget — our firmware development service builds this testing discipline into the project from day one, not as an afterthought. Tell us what your device has to do, and we will tell you how we will prove it does it. For the architecture side of the story, see our guide to embedded software architecture with RTOS; for getting hardware in hands fast enough to test against, see rapid hardware prototyping.
Contact us for customized solutions and quotes