A practical guide to WebRTC app development — when to self-host signaling, how embedded devices like ESP8266 join browser calls, and why Speex still beats Opus on tiny silicon. Lessons from four shipped projects.
Most teams begin a WebRTC app believing the hard part is the video call itself — getting two browsers to see and hear each other. That part is genuinely solved: the browser vendors did the heavy lifting, and a basic peer-to-peer call is a weekend project. The real engineering starts the moment the product requirements get serious: thousands of concurrent rooms, users on phones and desktops, and — the part that surprises most teams — dedicated hardware devices that need to join the same calls. At that point, WebRTC app development splits into three distinct problems: signaling, media on constrained endpoints, and codec selection. Get any one of them wrong and the product limps.
This guide walks through all three, grounded in four systems we actually built and shipped: an online interview platform with WebRTC video interviews across Android, iOS, and Web; a cloud video chat system with its own C++ media and signaling servers; an IP intercom running real-time voice on the decidedly modest ESP8266; and a smart helmet intercom system built around the F1C100S with Speex audio encoding. Each project forced a different tradeoff, and together they map the territory.
Here is the WebRTC standard's best-kept open secret: it does not define signaling at all. How peers discover each other, exchange session descriptions, and manage rooms is entirely up to you. That freedom is also a fork in the road — self-build your signaling infrastructure, or rent it from a CPaaS platform — and it is the single most consequential architectural decision in a WebRTC project.
Our online interview platform runs WebRTC video interviews across Android, iOS, and Web with fully self-built signaling. The backend combines Node.js and PHP, persists state in SQL Server, and implements its own room management — creating rooms, admitting participants, and tearing sessions down. Nothing about the interview workflow is generic: interviewers need waiting rooms, timed sessions, and role-based controls that a generic video API would have to be bent around. Owning the signaling layer meant the product logic and the call logic could evolve together instead of fighting through someone else's abstraction.
The cloud video chat system went further down the self-built path: both the media server and the signaling server are written in C++, the browser client needs no plugins, and multi-party video chat runs on infrastructure configured through a simple config.yml and deployed with an automated release.sh script. When you control the signaling server, features like custom authentication flows, fine-grained room permissions, and usage metering become straightforward engineering instead of integration puzzles.
Self-building is not machismo; it is arithmetic. It tends to win when:
Renting wins on exactly one axis, but it is an important one: time. If the product needs to demo next month and the team has no real-time experience, a managed platform buys calendar time that engineering cannot. The honest version of this decision is a bet on the product's future: rent when speed matters more than control, build when the product's core value lives inside the call itself. Both of our self-built projects started from the same judgment — the call workflow was the product — and neither team has regretted owning the layer.
WebRTC standardizes the media path and deliberately leaves signaling to you. That omission is not a gap in the spec — it is the decision that determines whether your product owns its future or rents it by the minute.
A desktop browser arrives at a WebRTC call carrying a full protocol stack: SDP offer/answer negotiation, ICE for NAT traversal, DTLS for handshake security, and SRTP for encrypted media. It has gigabytes of RAM and a multi-core CPU to throw at the problem. An ESP8266 has roughly 80 KB of usable RAM and a single modest core. It cannot run a full WebRTC stack — not slowly, not at all. This is the media negotiation gap, and bridging it is the second core problem of serious WebRTC development.
Our IP intercom project is the existence proof. It delivers real-time voice intercom on the low-cost ESP8266 — point-to-point and multi-party — using a self-developed audio transmission protocol rather than an off-the-shelf stack. The lesson is not that constrained devices can "do WebRTC" in the browser sense; it is that they can participate in real-time voice systems when the protocol is designed for their constraints: small frames, minimal handshake overhead, no cryptographic ceremony the chip cannot afford. Entry-level hardware can absolutely do real-time voice — it just cannot do it with browser-shaped protocols.
The smart helmet intercom system shows the production answer to the gap. The helmet's F1C100S device encodes audio with Speex and talks to a TCP audio relay server; a UDP heartbeat keeps the connection alive across flaky 4G links; group management lives in MySQL behind a PHP backend; and iOS and Android companion apps join the same conversations. The relay server is the Rosetta Stone: constrained devices speak a lightweight device protocol, richer clients speak whatever suits them, and the server translates, mixes, and routes.
This bridge pattern generalizes. In practice, almost every system that mixes browsers with embedded devices ends up with three tiers: browser-class endpoints running full WebRTC, device-class endpoints running a minimal purpose-built protocol, and a server tier that relays and transcodes between them. The key differences that force the split:
If your product roadmap includes any dedicated hardware — intercoms, helmets, door stations, wearables — plan for the bridge from day one. Bolting a gateway onto a browser-only architecture late in the project is where WebRTC timelines go to die.
Ask which audio codec a WebRTC app should use and the modern answer is automatic: Opus. It is the mandatory-to-implement codec in WebRTC browsers, it handles both voice and music gracefully, and its quality-per-bit is superb. On phones, desktops, and browsers, the discussion is over. On a small SoC, it is just beginning — which is why our helmet intercom encodes device-side audio with Speex, not Opus.
The F1C100S is a capable little application processor, but "capable" is relative: every CPU cycle spent on audio encoding is a cycle stolen from the application, and every kilobyte of codec state is memory that cannot hold sensor data or protocol buffers. Speex was designed in an era when these constraints were the norm. Its narrowband and wideband modes cover voice perfectly well, its CPU and RAM footprint are a fraction of Opus at comparable voice settings, and its fixed-point implementation runs efficiently on cores without generous floating-point hardware. For a helmet whose audio is human speech over 4G — not a concert stream — Speex delivers intelligible, low-latency voice at a cost the silicon can actually pay.
The practical rule: use Opus wherever the endpoint can afford it — browsers, phones, desktops — and reach for Speex (or a similarly light voice codec) on endpoints where the silicon budget says otherwise. Mixed systems transcode at the bridge: the relay server already exists to translate protocols, so translating codecs there too is a natural extension rather than a new problem.
Across these four projects, a handful of patterns kept reappearing wherever the systems met real users and real networks:
WebRTC app development rewards teams that respect its three separate problems: a signaling layer shaped around the product's actual workflow rather than a vendor's generic rooms, a bridge architecture that lets tiny devices and full browsers share the same conversations, and codec choices matched to each endpoint's silicon instead of copied from browser defaults. Our interview platform, cloud video chat, IP intercom, and helmet intercom each look different on the surface, but underneath they are the same three decisions made carefully — and made early, when they were still cheap.
If you are planning a real-time product — video interviews, multi-party conferencing, voice intercom on custom hardware, or anything that mixes browsers with embedded devices — our custom WebRTC app development services cover the full stack: self-hosted signaling design, embedded media integration down to ESP8266-class hardware, and codec and relay architecture tuned for your constraints. Tell us what your endpoints look like, and we will tell you what your architecture should be.
Contact us for customized solutions and quotes