embedded software

    Embedded Software Architecture: RTOS Task Design on Real Products

    How we structure firmware with FreeRTOS — task partitioning, real-time scheduling, and filesystem integration — based on three shipped embedded products.

    ·XTELL Engineering Team

    Why firmware architecture decides product reliability

    Most embedded products do not fail because one function was written badly. They fail because responsibilities were mixed together: control logic shares a task with logging, communication shares a buffer with the filesystem, and a temporary workaround becomes permanent. In the field, that kind of mixing shows up as frozen screens, missing alarm records, corrupted files, or devices that behave differently after a reboot. Good embedded software is rarely about clever code. It is about deciding, early, which task owns what.

    We have learned this across industrial controllers, security devices, and voice products. The pattern is consistent: when each task has one job, one owner, and a clear way to communicate with the rest of the system, the product becomes easier to test, easier to debug, and far more stable in real use. When tasks grow organically, reliability becomes accidental.

    This article explains how we approach that problem with FreeRTOS, using three shipped products as reference points. It is written for teams that already know the basics of embedded development and want a practical way to think about architecture. As an embedded development services team, we treat these choices as product decisions, not implementation details.

    Task partitioning with FreeRTOS: lessons from an STM32 PLC industrial control system

    In our STM32 PLC industrial control system, the hardware set the tone: an STM32F407 running FreeRTOS, FatFS for filesystem duties, SDIO for SD card storage, SRAM expansion for runtime memory, plus UART and GPIO/LED handling for industrial communication and status. The product needed dependable real-time scheduling, stable filesystem operation, and high-speed SDIO data recording. That combination is exactly where poor task design starts to hurt.

    Our approach was to let each layer do what it does best. The STM32F407 HAL drives the hardware. FreeRTOS manages multi-task scheduling. FatFS plus SDIO provides persistent storage. SRAM expands runtime memory. UART supports industrial communication. None of these responsibilities was allowed to leak into the others.

    Separate control, communication, and housekeeping

    The first decision was to separate timing domains. Control work belongs in tasks that are allowed to think only about control. Communication work belongs in tasks that are allowed to think only about moving bytes. Housekeeping—LEDs, GPIO status, diagnostics—belongs somewhere it cannot interfere with either.

    This sounds obvious, but it is easy to violate. A common mistake is to let a UART handler update application state directly, or to let a logging call block a control path while waiting for storage. We avoid both by making ownership explicit: a task owns its inputs, its outputs, and its failure behavior. If UART data needs to change control behavior, it travels through a defined interface rather than reaching across the system.

    The benefit is not only cleaner code. It changes debugging. When a timing problem appears, we know which task to inspect. When communication stalls, we know which interface to test. Architecture turns vague symptoms into local problems.

    Make storage boring: FatFS on SDIO

    Filesystem code is one of the most dangerous places to mix responsibilities. FatFS is powerful, but it should not be called casually from every task that feels like saving data. In the PLC system, storage had a single owner. Recording tasks prepared data; a dedicated storage task performed FatFS operations and wrote through SDIO to the SD card.

    This separation matters because storage is bursty and occasionally slow. Cards can be busy. Writes can take longer than expected. If those delays happen inside a control task, they become control problems. If they happen inside a storage task, they remain storage problems, and the rest of the system keeps running.

    We also treated recording as a pipeline rather than a side effect. Data entered through a clear boundary, was buffered in runtime memory supported by SRAM expansion, and was committed by the storage owner. That structure made high-speed SDIO data recording manageable without letting storage concerns contaminate scheduling.

    Stable storage is not a filesystem feature. It is an architecture feature: one owner, one path, and no surprise calls from unrelated tasks.

    Give every peripheral a home: UART, GPIO, LEDs, and SRAM

    UART, GPIO, and LED handling look trivial until they are scattered across the codebase. One task toggles an LED for status, another uses the same LED for errors, and a third repurposes a GPIO during a special mode. The result is hardware that lies to the operator.

    We give peripherals homes. UART handling lives with communication. GPIO and LED management lives with system status. Each has a small, explicit interface, so application tasks request behavior instead of manipulating pins. That may feel like extra structure on day one, but it prevents the slow drift that makes field diagnostics unreliable.

    SRAM expansion plays a supporting role here. It expands runtime memory for buffering and recording paths, but we do not use it as an excuse for careless ownership. Memory still has owners. Buffers still have producers and consumers. Expansion gives the architecture room to breathe; it does not replace the architecture.

    When the MCU is tiny: designing around constraints

    Not every product gets an STM32F407. Sometimes the constraint is the point: the device must be small, inexpensive, and good at one job. In those cases, architecture means refusing to ask a tiny MCU to do everything.

    STC32G128K: one small MCU, many 433MHz terminals

    Our 433MHz Security Linkage System used an STC32G128K running FreeRTOS to drive 433MHz terminals. It did not try to make that small MCU into a camera processor or a cloud gateway. The camera role belonged to a TXW82x SDK IPC/FPV camera, and status belonged to an LCD/OLED display. The embedded layer was deliberately split by capability.

    This is the key lesson for constrained designs: partition by subsystem, not just by task. FreeRTOS still helps organize the terminal work, but the larger architectural decision is which hardware owns which responsibility. A small MCU becomes reliable when its job is narrow and well defined.

    The same discipline applied to the surrounding system. A Flutter mobile app handled pairing, WiFi provisioning, live video, alarm push, and arming. The backend handled coordination. The embedded device stayed focused on terminals, camera integration, and local status. Nobody had to be heroic.

    ESP8266: real-time voice intercom without an off-the-shelf stack

    The IP intercom system pushed this idea further. On low-cost ESP8266 hardware, the team proved that real-time voice intercom could work on entry-level hardware, supporting both point-to-point and multi-party intercom. Instead of forcing a generic off-the-shelf stack onto the device, the project used an in-house audio transport protocol. The ESP8266 handled audio capture, encoding, and network transmission as one coherent path.

    This is not an argument that custom protocols are always better. It is an argument that architecture should follow the constraint. On entry-level hardware, every unnecessary layer has a cost. An in-house protocol kept the audio path under one design authority, so capture, encoding, and transmission could be reasoned about together.

    The broader lesson is to choose boring technology where the problem is ordinary, and custom technology where the constraint is real. Filesystem integration was boring on purpose. Voice transport on a tiny device was custom on purpose. Both decisions came from the same question: what does this product need to be reliable at?

    Think from firmware to cloud: one backend, three layers in sync

    Embedded architecture does not stop at the device edge. In the 433MHz system, the challenge was to keep 433MHz terminals, an IP camera, and a phone app in sync through one backend, with reliable live video and alarm push. That required thinking in three layers at once.

    The embedded layer stayed focused: STC32G128K with FreeRTOS managed the 433MHz terminals, the TXW82x SDK drove the camera, and LCD/OLED showed status. The mobile layer used a Flutter app for iOS and Android, with QR provisioning, RTSP/HLS video through VLC, alarm push, and arming controls. The backend used NestJS with PostgreSQL, Redis caching, EMQX for device messaging, ZLMediaKit for stream relay, and Docker Compose for deployment.

    The architecture worked because each layer had a clear contract. The backend handled JWT auth, caching, device messaging, and media relay. The app handled provisioning and user intent. The device handled terminals and camera duties. Live video and alarms were not afterthoughts bolted onto device code; they were paths with owners across all three layers.

    This is where many IoT projects drift. The device team invents its own message format, the app team invents another, and the backend becomes a translator. We prefer explicit contracts early: who authenticates, who publishes device messages, who relays streams, and how provisioning binds a device to an account. Those answers belong in the architecture, not in bug reports.

    A practical RTOS architecture checklist

    Use this as a review checklist before a design is considered stable. It reflects the patterns above, not a specific vendor or board.

    • Give each timing domain its own task: control, communication, storage, and housekeeping should not share execution paths.
    • Make storage single-owner: let one task own FatFS and SDIO access, with other tasks submitting work through a defined interface.
    • Separate control from communication: UART and network handling should not reach directly into control state.
    • Give peripherals explicit homes: UART, GPIO, LEDs, and displays each need one owner and one interface.
    • Use SRAM expansion deliberately: expand runtime memory for buffering and recording paths, but keep buffer ownership explicit.
    • Partition tiny systems by subsystem: let a small MCU own a narrow job, and move camera, UI, or gateway duties to hardware built for them.
    • Reserve custom protocols for real constraints: use an in-house transport only where entry-level hardware or real-time behavior demands it.
    • Keep device, app, and backend contracts explicit: authentication, device messaging, stream relay, provisioning, and alarm paths should all have owners.
    • Design failure paths early: storage busy, link loss, reboot, and reprovisioning are architecture inputs, not surprises.
    • Review task boundaries when features are added: every new feature should land in an existing owner or create a new one deliberately.

    Conclusion: architecture is a product feature

    Across these products, the same principle kept appearing. The STM32 PLC system stayed reliable because scheduling, storage, and peripherals had owners. The 433MHz system stayed coherent because a small MCU was given a narrow job and the backend, app, and device each had contracts. The ESP8266 intercom stayed usable because the audio path was designed as one path, not assembled from convenient parts.

    That is why firmware architecture deserves the same attention as hardware selection. Customers do not experience task diagrams, but they experience the results: devices that keep recording, alarms that arrive, voice that stays intelligible, and systems that recover cleanly. If your team is wrestling with tangled firmware, unstable storage, or a device that has outgrown its original structure, our firmware development services can help turn that architecture into a product advantage.

    Need Professional Services?

    Contact us for customized solutions and quotes