Topic

Speech Recognition

All digests tagged Speech Recognition

Building the Voice OS at Willow thumbnail

· 33:14

Building the Voice OS at Willow

This session details the engineering journey of building a Voice OS (Willow Voice) and expanding into advanced AI communication tools (Willow Scribe). The core challenge shifts from achieving accurate Speech Recognition (ASR) to making the output fast, reliable, and highly personalized. Key technical lessons include the necessity of low latency for user adoption, the use of LLMs for structuring unstructured human output (e.g., emails, Slack messages), and employing advanced techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to infer user intent and style while maintaining user privacy.

Key takeaways

  1. Latency is a primary adoption hurdle. 22:30

    While quantitative speed (e.g., typing vs. dictation) is important, the perceived speed (low latency) is critical. The initial delay in the Willow demo (3-5 seconds) caused users to revert to typing, demonstrating that immediate feedback is a top priority for product adoption. (Timestamp: 13:50)

  2. AI communication requires inferring style and tone. 27:10

    The challenge is moving beyond mere transcription to structuring the output (e.g., formal emails vs. casual Slack messages). This requires the model to infer the appropriate 'semi-natural language' (semi-linguistic style), which is highly variable based on context and user demographics. (Timestamp: 16:30)

  3. Privacy-preserving personalization is achievable.

    Personalization can be achieved without accessing semantic content by tracking limited, objective signals on the client side. By calculating metrics like the word-level Levenshtein distance (insertions, deletions, replacements) between the original and edited text, the system can infer user intent (e.g., 'deletion suggests the model was too verbose'). (Timestamp: 21:50)

  4. The future of work involves automating coordination.

    The speaker predicts that the majority of engineering time will shift from execution to coordination, defining requirements, and communicating. Future tools must act as sophisticated executive assistants, capable of understanding context and coordinating actions across people and agents. (Timestamp: 26:30)

Watch on YouTube Full article

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow thumbnail

· 15:05

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

Midam Kim presents a linguistic framework for diagnosing failures in voice AI, arguing that these failures are not isolated bugs but structured issues. She proposes that human communication is a 'joint activity' involving the continuous updating of a 'mental model.' The framework maps this process onto two channels (listening and speaking) and four interdependent levels: sounds, words, interaction, and mental model. Successful voice AI requires holistic orchestration across all these layers, rather than optimizing components (like ASR or TTS) in isolation.

Key takeaways

  1. Voice AI is a Joint Activity 5:00

    Human communication is a joint activity where both parties contribute sounds and words, continuously updating a shared mental model. Voice AI systems must replicate this joint nature to be effective.

  2. The Linguistic Framework 11:54

    The system must be analyzed across two channels (listening/speaking) and four interdependent levels: sounds, words, interaction, and mental model. Failure in one area (e.g., STT failure at the sound level) impacts the entire system.

  3. Mental Model Accumulation

    Unlike text chat where history remains visible, in voice interactions, sounds and words vanish. The only persistent element that matters for user satisfaction is the user's accumulating mental model.

  4. System Adaptability is Key

    The system must be designed to be dynamic, adapting to context, emotion, and language change over the course of the call, rather than functioning as a static pipeline.

Watch on YouTube Full article

Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium thumbnail

· 19:14

Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium

The talk analyzes the evolution of voice agents, arguing that current real-time voice models are fundamentally half-duplex (either listening or speaking). The core technical challenge is achieving full-duplex communication, which involves modeling overlapping speech (like backchanneling). The speaker, Neil Zeghidour, proposes that the most viable path forward is a hybrid architecture: coupling a small, highly natural, full-duplex speech-to-speech (S2S) interface with a powerful, asynchronous background text LLM to handle all complex reasoning and tool calling. This approach mitigates the inherent trade-off where improving naturalness sacrifices intelligence.

Key takeaways

  1. Evolution of Voice Agents 0:10

    Voice agent technology has progressed through constrained, closed-ended dialogue (Siri, 2011) to open-ended conversational models (OpenAI Voice Mode), and finally to agentic systems capable of real actions (e.g., ordering food).

  2. The Full-Duplex Challenge 12:28

    Human conversation is full-duplex, allowing for overlapping speech and backchanneling (e.g., 'Mhm, yeah'). Current S2S models, even with low latency, are limited by fundamental turn-taking mechanisms, making them feel unnatural.

  3. The Intelligence vs. Naturalness Trade-off 18:20

    There is a fundamental tension: every gain in naturalness (e.g., moving from cascaded STT/LLM/TTS to S2S) requires dedicating model capacity (weights) to audio modalities, which reduces the model's overall intelligence and reasoning capability.

  4. The Hybrid Solution 18:40

    The recommended approach is to split the system: use a small, on-device, full-duplex S2S model for natural conversation flow, while delegating all complex reasoning, tool calling, and agentic capabilities to a separate, powerful background text LLM.

Watch on YouTube Full article

Build voice-first apps with Gemini 3.5 Transcribe thumbnail

· 1:44

Build voice-first apps with Gemini 3.5 Transcribe

Google launched Gemini 3.5 Transcribe, an advanced LLM-based model designed for building voice-first applications. This model is available via both the Interactions API and the Live API, offering fast, contextually accurate transcription of multi-speaker recordings. Key strengths include superior recognition of structured data like email addresses and phone numbers, as well as robust support for over 70 different languages.

Key takeaways

  1. Model Availability

    Gemini 3.5 Transcribe is available on both the Interactions API and the Live API.

  2. Structured Data Recognition

    The LLM-based model excels at transcribing alphanumerics, such as email addresses (e.g., thorwebdev@google.com) and recognizing correct US phone number formats.

  3. Multi-Language Support

    The model can recognize and transcribe over 70 different languages, even when language hints are set to English.

Watch on YouTube Full article

How to build with Gemini 3.5 Transcribe thumbnail

· 4:50

How to build with Gemini 3.5 Transcribe

Google DeepMind launched Gemini 3.5 Transcribe, an LLM-based transcription model available via both the Interactions API and the Live API. This model significantly enhances accuracy by correctly transcribing complex data types—such as email addresses, phone numbers, and mixed units of measurement—and maintaining high performance across over 85 supported languages, even when language codes are set to English.

Key takeaways

  1. LLM-Based Transcription Model

    The model's LLM foundation allows it to handle complex data structures and context better than traditional transcription models. For example, it can correctly identify and edit email addresses even if spoken phonetically (e.g., 'tosten at google.com').

  2. Handling Complex Data Types 2:00

    Gemini 3.5 Transcribe accurately recognizes specific formats, including US phone numbers and international variations (e.g., Singapore's 8-digit format). It can also correctly interpret units of measure (e.g., meters vs. centimeters).

  3. Multilingual and Customization Support 0:40

    The model supports over 85 languages, automatically recognizing spoken language even if language hints are set to English. Accuracy can be further improved by providing custom vocabulary (e.g., names of people in a meeting) or setting specific language codes.

Watch on YouTube Full article