AIoT engineering guide — cloud vs edge inference tradeoffs, TinyML on ESP32-S3, speech recognition on microcontrollers, and vision AI at the edge. Lessons from pose detection, smart BBQ vision, and voice interaction systems we shipped.
Every AIoT project starts with an exciting demo: a camera that recognizes poses, a grill that identifies food, a box that understands speech. Then comes the question teams postpone until it hurts — where does the inference actually run? On a cloud GPU or on the device itself? The answer shapes the hardware cost, the latency, the privacy story, and whether the product works at all when the WiFi drops.
We have built AI into embedded products across the spectrum: cloud-assisted vision systems, on-device neural networks, and microcontroller speech recognition. This guide lays out how to make the cloud-vs-edge decision, what TinyML can realistically do on a microcontroller in 2026, and the engineering patterns that survive contact with real users.
The naive view is that cloud inference is "easy" and edge inference is "hard." The honest view is a tradeoff table with no universally right answer:
TinyML — machine learning on microcontrollers with kilobytes of RAM — has crossed from research curiosity to shipping technology. The ESP32-S3 speech recognition system we built (ESP32-S3 voice box) is a working proof: using the ESP-IDF framework and the ESP-SR speech library, it does voice wake-up, audio streaming over WebSocket, and voice interaction with LCD display and MP3 playback — on a microcontroller. No cloud needed for the wake word; the device listens, recognizes, and only then talks to the network.
What TinyML does well in 2026:
What it does not do: open-ended understanding. Our AI Home voice assistant runs the full STT → LLM agent → TTS pipeline with the language model on a server, because a microcontroller cannot hold a conversation. The honest architecture is hybrid — the edge handles the always-on, latency-critical, privacy-sensitive part, and the cloud handles the part that needs a data center.
The smart BBQ vision system recognizes 20+ common food ingredients, personalizes seasoning, and optimizes cooking — improving cooking accuracy by 85% — with remote monitoring and automatic cleaning. The vision model runs close to the device because a grill controller cannot wait for a cloud round-trip to decide the steak is done, and because nobody wants their backyard barbecue streamed to a data center.
The pose detection system pushes further: 17 keypoints, 30 FPS, 8 simultaneous people, 95%+ accuracy, running on mobile-class hardware at low power. The engineering lesson from both projects is the same — edge vision AI is a systems problem, not just a model problem. The model is maybe a third of the work; the rest is the camera pipeline, the frame buffering, the thermal management (neural accelerators get hot), and the fallback behavior when confidence is low.
Across our projects, the architecture that survives production most often is a three-tier split:
The tiers communicate in events, not streams: the device sends "person detected at the door" rather than streaming video 24/7. This pattern cuts bandwidth by orders of magnitude and is the reason the products stay responsive and affordable.
Picking the AIoT chip is where many projects go wrong, usually by over- or under-buying compute:
For all the momentum behind edge AI, the cloud remains the right answer in three situations. First, when the model changes faster than your firmware release cycle — recommendation models, fraud detection, anything retrained weekly belongs behind an API. Second, when the workload needs resources no edge device will have soon: large language models, video generation, or training on fleet data. Our NLP service project (GPT-2 based web inference) lives in the cloud for exactly this reason. Third, when the device is already cloud-connected for other reasons and the inference is infrequent — a daily image classification for a dashboard does not justify edge silicon.
The mature AIoT team does not pick a side; it picks per workload. Latency-critical, privacy-sensitive, always-on inference goes to the edge. Fast-moving, heavyweight, or fleet-aggregated intelligence stays in the cloud. Revisit the split every product generation — the boundary moves as silicon gets cheaper and models get smaller.
A model trained on a workstation rarely runs on a microcontroller unchanged. The bridge is optimization, and the two techniques that matter most are quantization and pruning. Quantization reduces the numerical precision of the model's weights — from 32-bit floats to 8-bit integers — typically shrinking the model to a quarter of its size with a 1–2% accuracy cost. For the keyword-spotting and simple classification workloads that dominate TinyML, that tradeoff is overwhelmingly worth it. Pruning goes further by removing the neural connections that contribute least, producing a smaller, faster model that does the same job.
The workflow that works in practice: train in full precision, quantize, then measure on the actual target — not in a simulator. The ESP32-S3 speech recognition pipeline went through exactly this loop, with each iteration benchmarked for wake-word accuracy against false-trigger rate in noisy rooms. A model that is 98% accurate in the lab and triggers on every television commercial is not a 98%-accurate model. Budget time for collecting real-world test audio or images from the deployment environment; the gap between lab data and field data is where edge AI projects most often stumble.
AIoT is not about putting the biggest model on the smallest chip — it is about putting the right intelligence in the right tier. TinyML handles the always-on simple stuff, edge SoCs handle real-time vision, and the cloud handles language and fleet learning. The projects that ship are the ones that made the tier decision deliberately, benchmarked on real hardware, and designed the model update path before the first unit left the factory.
If you are planning an AI-powered device — vision, voice, or sensor intelligence — our IoT development service covers the full stack from model selection and TinyML porting to cloud training pipelines. For the system architecture view, see our IoT system architecture guide; for the on-device UI side, see LVGL GUI development on MCUs.
Contact us for customized solutions and quotes