AI Engineer

Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium

Published 2026-09-15 · Duration 19:14

Summary

The talk analyzes the evolution of voice agents, arguing that current real-time voice models are fundamentally half-duplex (either listening or speaking). The core technical challenge is achieving full-duplex communication, which involves modeling overlapping speech (like backchanneling). The speaker, Neil Zeghidour, proposes that the most viable path forward is a hybrid architecture: coupling a small, highly natural, full-duplex speech-to-speech (S2S) interface with a powerful, asynchronous background text LLM to handle all complex reasoning and tool calling. This approach mitigates the inherent trade-off where improving naturalness sacrifices intelligence.

Download summary

Key takeaways

  1. Evolution of Voice Agents 0:10

    Voice agent technology has progressed through constrained, closed-ended dialogue (Siri, 2011) to open-ended conversational models (OpenAI Voice Mode), and finally to agentic systems capable of real actions (e.g., ordering food).

  2. The Full-Duplex Challenge 12:28

    Human conversation is full-duplex, allowing for overlapping speech and backchanneling (e.g., 'Mhm, yeah'). Current S2S models, even with low latency, are limited by fundamental turn-taking mechanisms, making them feel unnatural.

  3. The Intelligence vs. Naturalness Trade-off 18:20

    There is a fundamental tension: every gain in naturalness (e.g., moving from cascaded STT/LLM/TTS to S2S) requires dedicating model capacity (weights) to audio modalities, which reduces the model's overall intelligence and reasoning capability.

  4. The Hybrid Solution 18:40

    The recommended approach is to split the system: use a small, on-device, full-duplex S2S model for natural conversation flow, while delegating all complex reasoning, tool calling, and agentic capabilities to a separate, powerful background text LLM.

Technical details

  • Audio Tokenization and Compression 890s

    Training LLMs on raw audio (a waveform) is computationally prohibitive. For example, 8 words can generate 72,000 time steps at 24 kHz. This is solved by using neural codecs (audio tokenizers) that compress raw audio into a dense, token-like representation, allowing the LLM to process the audio domain abstractly.

  • Multi-Stream Language Models 970s

    To achieve full-duplex capability, the model must move beyond single-sequence prediction (half-duplex) and model two independent token streams simultaneously. This requires multi-stream language models.

  • System Architecture Comparison 150s

    The pipeline has evolved from a constrained STT $ ightarrow$ NLU $ ightarrow$ TTS (Siri) to a single, end-to-end S2S model (absorbing all steps), and finally to a hybrid architecture (S2S interface + background text LLM).

Mentioned resources

  • Gradium (Company/Model Developer)
  • Moshi (Full Duplex S2S Model)
  • Hibiki (Real-time S2S Translation System)
  • GPT-4o (Text LLM/Model)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.