Topic

Real-Time Systems

All digests tagged Real-Time Systems

The next generation of voice AI with Google DeepMind and Sierra AI thumbnail

· 5:22

The next generation of voice AI with Google DeepMind and Sierra AI

The discussion outlines the evolution of voice AI from traditional pipelines to advanced native audio models, focusing on achieving truly real-time, conversational experiences. Key advancements include offering specialized models (lightweight for speed, enterprise for precision), improving metrics beyond Word Error Rate (WER) to measure conversational flow, and enabling seamless multilingual code-switching and complex, multi-step agentic tasks.

Key takeaways

  1. Dual Model Architecture

    Developers now have two options: a lightweight, faster model for quick conversations, and a more robust, enterprise-grade model designed for high-stakes environments requiring multi-step function calling and high accuracy.

  2. Advanced Latency Metrics 2:28

    Conversational quality is measured by two critical latencies: Time to First Audio (TFA) and Time to First Useful Response (TFUR), both of which must be minimized to maintain a natural, uninterrupted dialogue flow.

  3. Multilingual Code-Switching

    Native audio models are highly effective at understanding and navigating language shifts and mixed-language phrasing, avoiding the 'broken telephone' effect common in traditional text-based transcription setups.

Watch on YouTube Full article

Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI thumbnail

· 13:14

Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI

This presentation details the architecture of a real-time voice agent harness designed to achieve Frontier-level intelligence while maintaining low latency. The system utilizes a cascaded voice stack, drawing parallels to self-driving car systems, breaking the process into Perception (Transcription), Planning (LLM/Tool Calling), and Control (Speech Synthesis). Key innovations include a streaming speculative transcriber for accuracy, background agents for tool calling, and a prefix cache combined with audio suppression techniques to hide generation latency and ensure seamless, natural conversation flow.

Key takeaways

  1. Cascaded Voice Agent Architecture

    The system is structured into three layers: Perception (Transcription, converting audio to data), Planning (LLM, processing data and determining actions), and Control (Speech Synthesis, converting text back to natural audio).

  2. Streaming Speculative Transcriber 2:32

    A hybrid approach combining a fast streaming transcriber (e.g., Flux) with a slower, more accurate batch transcription (e.g., Scribe V2) that uses context to correctly identify entities like names and dates of birth.

  3. Background Tool Calling 7:04

    To reduce round trips with slow, intelligent LLMs, background agents perform tool calls and inject the results into the main model's context, making the main agent believe it executed the call itself.

  4. Prefix Cache for Synthesis 10:57

    The prefix cache monitors the model's output stream, checking if audio for a sequence of words already exists from a prior turn. This allows the agent to start speaking immediately from cached audio while the rest of the sentence is being generated.

Watch on YouTube Full article