AssemblyAI

Universal 3.5 Pro Demo: Smarter Speech-to-Text with Contextual Awareness

Published 2026-07-01 · Duration 10:07

Summary

This demo introduces Universal 3.5 Pro, an advanced Speech-to-Text (STT) model designed to significantly boost transcription accuracy through enhanced contextual awareness. Key features include passing domain-specific prompts (e.g., 'cardiology consultation'), applying context to key terms to prevent misapplication, and supporting dynamic mid-call prompt updates via API calls. Furthermore, the model retains conversation history (agent context), allowing it to accurately transcribe user input even in poor audio conditions by understanding the situational flow of a voice agent interaction.

Download summary

Key takeaways

  1. Contextual Prompting for Domain Accuracy

    Passing detailed information about the audio content (e.g., 'cardiology consultation between Dr. Smith and elderly patient regarding chest pain...') dramatically improves model accuracy within specific domains. The more specific the prompt, the better the results.

  2. Contextual Key Terms 2:00

    Unlike previous methods where key terms were applied blindly, Universal 3.5 Pro allows users to define what a key term represents (e.g., 'The user's name is Zachary Klebanoff'). This prevents the model from incorrectly applying terminology based solely on acoustic similarity.

  3. Dynamic Mid-Call Prompt Updates 2:55

    The prompt can be updated in real time via the API (not available in the playground demo). This is crucial for voice agents, allowing tool calls or external data to adjust the model's context mid-conversation.

  4. Conversation/Agent Context 3:30

    The model retains previous transcriptions and accepts LLM-generated responses from a voice agent as context. This provides situational awareness, improving accuracy even in poor audio conditions and reducing the Word Error Rate (WER) on voice agent datasets.

Technical details

  • Contextual Prompting 45s

    The model accepts descriptive prompts to guide transcription, ranging from general domains (e.g., 'medical consultation call') to highly specific scenarios.

  • Key Term Context Application 120s

    Contextual prompting allows the model to distinguish between acoustically similar phrases and defined key terms, ensuring accurate transcription even when ambiguity exists (e.g., differentiating a name from common speech).

  • API Integration for Dynamic Context 175s

    Mid-call prompt updates are achieved through the API, enabling systems to inject new contextual information (e.g., 'the caller's name is Lance Armstrong and his bike tire is popped') as events occur.

  • Agent Context Flow 210s

    The system automatically retains speech-to-text transcripts within the session. Developers can pass LLM-generated agent responses directly to the model, providing a complete conversational history for improved accuracy.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.