Building Low-Latency Two-Way Voice Agents with Gemini Multimodal Live WebSocket Audio API in Python

Building Low-Latency Two-Way Voice Agents with Gemini Multimodal Live WebSocket Audio API in Python

(Updated: ) 📖 1 min read

Traditional AI voice agents rely on a brittle three-step cascade:

  1. Speech-to-Text (STT) (e.g. Whisper) ➔ 800ms
  2. Text LLM Reasoning (e.g. GPT-4o) ➔ 1200ms
  3. Text-to-Speech (TTS) (e.g. ElevenLabs) ➔ 700ms Total Turn Latency: 2.7 to 3.5 seconds — far too slow for natural human conversation.

Google’s Gemini Multimodal Live API solves this by accepting direct bidirectional raw audio streams over WebSockets, cutting response onset to sub-400ms.


1. Architectural Blueprint

┌──────────────┐      Raw PCM 16kHz Audio Stream      ┌──────────────────────┐
│ Microphone / │ ───────────────────────────────────> │ Gemini Live          │
│ WebRTC Client│ <─────────────────────────────────── │ WebSocket Gateway    │
└──────────────┘      Streamed 24kHz Audio Output     │ (Native Multimodal)  │
                                                      └──────────────────────┘

2. Python Asynchronous Client Implementation

import asyncio
import os
import json
import websockets

GEMINI_WS_URL = "wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent"

async def run_voice_agent():
    api_key = os.environ["GEMINI_API_KEY"]
    url = f"{GEMINI_WS_URL}?key={api_key}"

    async with websockets.connect(url) as ws:
        # 1. Send Handshake Setup Configuration
        setup_msg = {
            "setup": {
                "model": "models/gemini-2.0-flash-exp",
                "generation_config": {
                    "response_modalities": ["AUDIO"],
                    "speech_config": {
                        "voice_config": {
                            "prebuilt_voice_config": {"voice_name": "Puck"}
                        }
                    }
                }
            }
        }
        await ws.send(json.dumps(setup_msg))
        setup_resp = await ws.recv()
        print("✅ Session Initialized:", setup_resp)

        # 2. Producer/Consumer Tasks for Bidirectional Streaming
        async def send_audio():
            # In production: read from PyAudio or WebRTC audio buffer
            while True:
                dummy_pcm_chunk = b"\x00" * 3200 # 100ms of 16kHz 16-bit audio
                payload = {
                    "realtime_input": {
                        "media_chunks": [
                            {"mime_type": "audio/pcm;rate=16000", "data": dummy_pcm_chunk.hex()}
                        ]
                    }
                }
                await ws.send(json.dumps(payload))
                await asyncio.sleep(0.1)

        async def receive_audio():
            while True:
                msg = await ws.recv()
                data = json.loads(msg)
                # Extract and play server audio stream chunks
                if "serverContent" in data:
                    print("🔊 Received audio token frame from Gemini Live")

        await asyncio.gather(send_audio(), receive_audio())

if __name__ == "__main__":
    asyncio.run(run_voice_agent())

3. Production Best Practices

  • Barge-In Handling: Immediately terminate local speaker buffers when serverContent.interrupted is flagged.
  • Noise Gating: Enforce client-side Voice Activity Detection (VAD) via WebRTC VAD or Silero VAD to avoid streaming silent background room noise.
  • WebSocket Reconnections: Implement exponential backoff reconnects to maintain conversational continuity during transient network glitches.
FREE BOILERPLATE

Get the Gemini Live WebSocket Audio Starter Boilerplate

A production-grade Voice AI FastAPI backend and real-time HTML/JS sandbox template to jumpstart your audio agent projects.

Professor XAI
Professor XAI ML Engineer passionate about advancing AI technologies and building intelligent systems.
comments powered by Disqus