Workshop: Async, Sync, or Realtime: Which model do I choose?
Summary
This webinar provides a comprehensive framework for selecting the appropriate Speech-to-Text (STT) model—Async, Sync, Realtime, or Dictation—based on a product's specific requirements for latency, accuracy, and cost. The choice hinges on whether the audio is pre-recorded (Async) or if a user is actively waiting for the text (Realtime/Sync). For advanced applications like voice agents, specialized models such as Universal 3.6 Pro are recommended for handling turn-taking context and background noise. Developers are cautioned against common pitfalls, such as using Async for live experiences or arbitrarily chunking audio, which can severely degrade transcription quality.
Key takeaways
-
API Selection Framework
0:02
The core decision point is whether a user is waiting for the text in real-time. If no one is waiting, Async is suitable. If someone is waiting, Realtime or Sync must be considered.
-
Voice Agent Requirements
0:23
For voice agents (where the system talks back), use the state-of-the-art Universal 3.6 Pro model. This model handles turn-passing context, conversation memory, and background noise filtering.
-
Dictation vs. Realtime
0:45
Dictation is designed for short clips and provides both the verbatim transcript and an LLM-cleaned-up version in a single request, simplifying the process compared to multiple requests.
-
Avoiding Chunking Pitfalls
0:14
When processing audio, do not chunk arbitrarily (e.g., 30-second clips). Always chunk based on natural utterances (spoken phrases) to maintain high word error rate accuracy.
Technical details
-
API Workflow Comparison
8s
Async processes large, pre-recorded files (up to 10 hours) via webhooks/pull. Sync handles short, chopped clips (max 2 minutes) with a quick request/response. Realtime streams words via a Websocket connection. Dictation is a single request providing both verbatim and cleaned text.
-
Model Differentiation (Pro vs. Base)
30s
The Universal 3.6 Pro model is for complex, interactive scenarios (voice agents) requiring advanced features like turn detection and context passing. The Universal 3.6 Base model is simpler, cheaper, and optimized for pure transcription tasks like notetaking or live captions.
-
Latency Benchmarks
57s
Sync and Dictation offer near-instantaneous round-trip latency (approx. 1 second). Realtime provides initial words within 1 second of connection, with final text taking 3-4 seconds after audio ends. Async latency is typically 10 to 32 seconds for a full file.
-
Advanced Features and Pitfalls
120s
Speaker diarization (identifying who spoke what) is available as an add-on feature for Async. Pitfalls include: 1) Using Async when real-time interaction is needed. 2) Sending pre-recorded files through the Realtime stream (use Async instead). 3) Arbitrarily chunking audio instead of using a Voice Activity Detector (VAD) and sending to Sync.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.