Building conversational voice AI that feels genuinely human is an uncompromising game of latency budgeting. When two people speak, natural conversational turn-taking happens in roughly 220ms to 280ms. The moment an AI agent takes longer than 450ms to formulate an audio response, the psychological illusion of conversation shatters, turning the interaction into an awkward walkie-talkie exchange.
In this architecture breakdown, we dissect the complete end-to-end engineering pipeline required to achieve sub-300ms total glass-to-glass latency in production voice agents.
1. The 300ms Latency Budget Breakdown
Every millisecond must be accounted for across the network, processing, and generation layers:
| Pipeline Stage | Technology Stack | Allotted Budget (Target) | Failure / Slow Path |
|---|---|---|---|
| Network Ingress | WebRTC OPUS Stream | 25ms - 40ms | >120ms (WebSockets TCP retransmits) |
| Voice Activity Detection (VAD) | Silero VAD / Edge ONNX | 15ms - 25ms | >80ms (Server-side chunk buffering) |
| Speech-to-Text (STT) | Streaming Whisper / Conformer | 50ms - 80ms (First Chunk) | >300ms (Batch audio file uploads) |
| LLM Time-to-First-Token (TTFT) | Frontier Distilled Model / vLLM | 60ms - 90ms | >400ms (Un-cached generic API calls) |
| Text-to-Speech (TTS) Synthesis | Streaming Neural Vocoder | 40ms - 60ms (First Wave) | >250ms (Full-sentence synthesis) |
| Audio Egress & Playback | WebRTC SRTP Playback | 20ms - 30ms | >80ms (Jitter buffer under-runs) |
| Total Glass-to-Glass | Full Pipeline | 210ms - 325ms | >1,200ms (Standard naive stacks) |
2. Why WebRTC Trumps WebSockets
Most first-generation voice bots piped raw PCM audio over persistent WebSockets. Over unstable 4G or 5G mobile connections, WebSockets suffer from TCP Head-of-Line Blocking: if a single audio packet drops, the entire transmission halts until the missing byte is retransmitted, causing audible micro-stutters and delayed responses.
WebRTC uses SRTP over UDP. If an audio frame drops, the jitter buffer interpolates the loss, maintaining stream cadence without blocking subsequent chunks.
┌────────────────────────────────────────────────────────┐
│ SUB-300MS REAL-TIME VOICE PIPELINE │
└────────────────────────────────────────────────────────┘
│
┌───────────────────────────────┬───────┴───────────────────────┬───────────────────────────────┐
▼ ▼ ▼ ▼
┌──────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ WebRTC Audio │ (30ms) │ Streaming STT │ (70ms) │ LLM Speculative │ (80ms) │ Streaming TTS │
│ Ingress ├─────────────►│ Transcribe VAD ├────────────►│ First Token ├────────────►│ Audio Egress │
└──────────────┘ └─────────────────┘ └─────────────────┘ └─────────────────┘
3. Seamless Barge-In (Interruption) Mechanics
The most challenging engineering hurdle in voice agents is handling interruptions smoothly. When a user interrupts while the assistant is speaking:
- Local VAD Detection: The edge client detects speech within 60ms.
- Immediate Audio Ducking: Mute local playback volume to 0 within 5ms.
- Cancellation Signal: Emit an out-of-band SCTP datachannel cancellation event to the backend.
- Server Drain: The inference server aborts the LLM token generation queue and clears the TTS synthesiser buffer.
4. Production Checklist for Voice AI Engineers
- Deploy Audio Processing Near the User: Place WebRTC media servers within 30ms network ping of your primary target audience.
- Use Sentence Chunking with Punctuation Lookahead: Feed the first 4-word phrase directly to the TTS engine so speech playback starts while the LLM is still drafting the conclusion.
- Never Synthesize Silence: Keep acoustic room tone humming softly so users never feel like the line dropped.
Get the Gemini Live WebSocket Audio Starter Boilerplate
A production-grade Voice AI FastAPI backend and real-time HTML/JS sandbox template to jumpstart your audio agent projects.