Weights & Biases

Building the Voice OS at Willow

Published 2026-10-07 · Duration 33:14

Summary

This session details the engineering journey of building a Voice OS (Willow Voice) and expanding into advanced AI communication tools (Willow Scribe). The core challenge shifts from achieving accurate Speech Recognition (ASR) to making the output fast, reliable, and highly personalized. Key technical lessons include the necessity of low latency for user adoption, the use of LLMs for structuring unstructured human output (e.g., emails, Slack messages), and employing advanced techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to infer user intent and style while maintaining user privacy.

Download summary

Key takeaways

  1. Latency is a primary adoption hurdle. 22:30

    While quantitative speed (e.g., typing vs. dictation) is important, the perceived speed (low latency) is critical. The initial delay in the Willow demo (3-5 seconds) caused users to revert to typing, demonstrating that immediate feedback is a top priority for product adoption. (Timestamp: 13:50)

  2. AI communication requires inferring style and tone. 27:10

    The challenge is moving beyond mere transcription to structuring the output (e.g., formal emails vs. casual Slack messages). This requires the model to infer the appropriate 'semi-natural language' (semi-linguistic style), which is highly variable based on context and user demographics. (Timestamp: 16:30)

  3. Privacy-preserving personalization is achievable.

    Personalization can be achieved without accessing semantic content by tracking limited, objective signals on the client side. By calculating metrics like the word-level Levenshtein distance (insertions, deletions, replacements) between the original and edited text, the system can infer user intent (e.g., 'deletion suggests the model was too verbose'). (Timestamp: 21:50)

  4. The future of work involves automating coordination.

    The speaker predicts that the majority of engineering time will shift from execution to coordination, defining requirements, and communicating. Future tools must act as sophisticated executive assistants, capable of understanding context and coordinating actions across people and agents. (Timestamp: 26:30)

Technical details

  • Speech Recognition & ASR Improvement 730s

    The adoption of voice AI was accelerated by the improved accuracy of ASR models (e.g., Whisper Large V3 Turbo, Speechmatics Deep, Gram, AssemblyAI). This reduced the Word Error Rate (WER) to a minimum, making dictation viable in noisy environments and eliminating the need for constant manual correction. (Timestamp: 7:30)

  • LLM Post-Processing for ASR 1210s

    Raw ASR output is often unstructured and grammatically messy. To make it usable, the output must be processed by an LLM to impose structure and readability, moving the output from a 'stream of consciousness' to a polished, deliverable format. (Timestamp: 12:10)

  • Model Scaling and Tuning 1430s

    The speaker tested using a smaller, efficient model (Llama 3 1 8B) and attempted Prompt Tuning to achieve the quality of a frontier model (4O). This demonstrated that while scaling down is necessary for speed, achieving high quality requires careful tuning and, often, Supervised Fine-Tuning (SFT). (Timestamp: 14:30)

  • Reinforcement Learning (RL) for Alignment 1830s

    RL was used to refine the model's behavior by focusing on objective, measurable failure patterns (e.g., confusing 'pounds as weight' vs. 'pounds as currency') rather than subjective human evaluation. This required creating a structured data classification studio. (Timestamp: 18:30)

  • Few-Shot Learning and Style Profiling

    To achieve personalization, the system must be fed multiple examples (few-shot) of the user's writing style (e.g., use of exclamation marks, capitalization, or specific phrasing) to create a 'style profile' that guides the LLM's output. (Timestamp: 22:30)

Mentioned resources

  • Willow Voice (Product)
  • Code X (Product/Platform)
  • Cloud Code (Product/Platform)
  • Whisper Large V3 Turbo (ASR Model)
  • Speechmatics Deep (ASR Model)
  • Gram (ASR Model)
  • AssemblyAI (ASR Model)
  • ElevenLabs (ASR Model)
  • Llama 3 1 8B (LLM Model)
  • Willow Scribe (Product)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.