Google Developers

The next generation of voice AI with Google DeepMind and Sierra AI

Published 2026-09-24 · Duration 5:22

Summary

The discussion outlines the evolution of voice AI from traditional pipelines to advanced native audio models, focusing on achieving truly real-time, conversational experiences. Key advancements include offering specialized models (lightweight for speed, enterprise for precision), improving metrics beyond Word Error Rate (WER) to measure conversational flow, and enabling seamless multilingual code-switching and complex, multi-step agentic tasks.

Download summary

Key takeaways

  1. Dual Model Architecture

    Developers now have two options: a lightweight, faster model for quick conversations, and a more robust, enterprise-grade model designed for high-stakes environments requiring multi-step function calling and high accuracy.

  2. Advanced Latency Metrics 2:28

    Conversational quality is measured by two critical latencies: Time to First Audio (TFA) and Time to First Useful Response (TFUR), both of which must be minimized to maintain a natural, uninterrupted dialogue flow.

  3. Multilingual Code-Switching

    Native audio models are highly effective at understanding and navigating language shifts and mixed-language phrasing, avoiding the 'broken telephone' effect common in traditional text-based transcription setups.

Technical details

  • Conversational Latency Benchmarking 148s

    The goal is to move beyond static context multi-turn benchmarks. Key metrics include minimizing Time to First Audio (TFA) and Time to First Useful Response (TFUR). The industry is moving toward user simulator benchmarks that incorporate real noise and diverse personas.

  • Agentic Voice Tasks 0s

    The use case for agents (e.g., outbound sales agents) requires long-horizon reasoning and complex orchestration, necessitating models capable of multi-step function calling and stateful task management while maintaining a natural vocal flow.

  • Voice Quality Hierarchy 240s

    Voice quality is viewed in a hierarchy: 1) Accuracy (Can the job be done?), 2) Quality (Is the call tolerable?), and 3) Experience (Is it pleasant and natural to listen to?).

Mentioned resources

  • Google DeepMind (Product/Company)
  • Sierra AI (Product/Company)
  • TAU (Benchmark)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.