Workshop: Async, Sync, or Realtime: Which model do I choose?
This webinar provides a comprehensive framework for selecting the appropriate Speech-to-Text (STT) model—Async, Sync, Realtime, or Dictation—based on a product's specific requirements for latency, accuracy, and cost. The choice hinges on whether the audio is pre-recorded (Async) or if a user is actively waiting for the text (Realtime/Sync). For advanced applications like voice agents, specialized models such as Universal 3.6 Pro are recommended for handling turn-taking context and background noise. Developers are cautioned against common pitfalls, such as using Async for live experiences or arbitrarily chunking audio, which can severely degrade transcription quality.
Key takeaways
-
API Selection Framework
0:02
The core decision point is whether a user is waiting for the text in real-time. If no one is waiting, Async is suitable. If someone is waiting, Realtime or Sync must be considered.
-
Voice Agent Requirements
0:23
For voice agents (where the system talks back), use the state-of-the-art Universal 3.6 Pro model. This model handles turn-passing context, conversation memory, and background noise filtering.
-
Dictation vs. Realtime
0:45
Dictation is designed for short clips and provides both the verbatim transcript and an LLM-cleaned-up version in a single request, simplifying the process compared to multiple requests.
-
Avoiding Chunking Pitfalls
0:14
When processing audio, do not chunk arbitrarily (e.g., 30-second clips). Always chunk based on natural utterances (spoken phrases) to maintain high word error rate accuracy.