Real-time is a product decision before it's a technical one
Notes from building WebRTC and WebSocket systems for interviews, therapy sessions, and multiplayer research tasks.
I’ve built real-time systems for three fairly different purposes: live technical interviews at MLPro, multiplayer research tasks for diagnostic studies, and now CHIME — a music co-creation platform where autistic adolescents and music therapists play together in a shared session.
They all use roughly the same primitives. They needed almost entirely different tradeoffs.
Latency budgets come from the human activity
An interview tolerates 200ms of audio latency. A conversation adapts. Musical co-creation does not — if two people are trying to play in time together, the round trip has to sit inside the window where the brain still perceives simultaneity, and that budget is brutal. Deciding what “real-time” means for your activity, in milliseconds, before choosing a transport, avoids a category of architecture you can’t refactor your way out of.
Failure modes are what you’re actually designing
The happy path is the same everywhere. What differs is what should happen when the network degrades:
- In an interview, drop video, keep audio, and keep the code editor authoritative. Losing the candidate’s face is survivable; losing their work is not.
- In a research session, the session state is data. A dropped participant mid-task may invalidate the trial, so reconnection has to restore exact state, and the system needs to record honestly that a degradation occurred.
- In a therapy session, the priority is the relationship. Preserve continuity and never make the participant feel they broke something.
Those are product decisions. They dictate the technical design, not the other way around.
Peer-to-peer isn’t automatically the answer
WebRTC’s mesh model is elegant for two participants and stops being elegant quickly. The moment you need server-side recording, more than a few peers, or a canonical record of what happened, you’re routing through a server anyway. Deciding that early is much cheaper than discovering it after you’ve built around the assumption that peers talk directly.
The unglamorous conclusion
Most of the hard work in real-time systems isn’t the media stack — it’s session state, presence, reconnection, and clock alignment. The media is a solved problem with good libraries. Knowing who is in the room, what they can see, and what happened when they disappeared is the part you’ll write yourself.