Topic

Voice Agent Development

All digests tagged Voice Agent Development

Workshop: Async, Sync, or Realtime: Which model do I choose? thumbnail

· 30:08

Workshop: Async, Sync, or Realtime: Which model do I choose?

This webinar provides a comprehensive framework for selecting the appropriate Speech-to-Text (STT) model—Async, Sync, Realtime, or Dictation—based on a product's specific requirements for latency, accuracy, and cost. The choice hinges on whether the audio is pre-recorded (Async) or if a user is actively waiting for the text (Realtime/Sync). For advanced applications like voice agents, specialized models such as Universal 3.6 Pro are recommended for handling turn-taking context and background noise. Developers are cautioned against common pitfalls, such as using Async for live experiences or arbitrarily chunking audio, which can severely degrade transcription quality.

Key takeaways

  1. API Selection Framework 0:02

    The core decision point is whether a user is waiting for the text in real-time. If no one is waiting, Async is suitable. If someone is waiting, Realtime or Sync must be considered.

  2. Voice Agent Requirements 0:23

    For voice agents (where the system talks back), use the state-of-the-art Universal 3.6 Pro model. This model handles turn-passing context, conversation memory, and background noise filtering.

  3. Dictation vs. Realtime 0:45

    Dictation is designed for short clips and provides both the verbatim transcript and an LLM-cleaned-up version in a single request, simplifying the process compared to multiple requests.

  4. Avoiding Chunking Pitfalls 0:14

    When processing audio, do not chunk arbitrarily (e.g., 30-second clips). Always chunk based on natural utterances (spoken phrases) to maintain high word error rate accuracy.

Watch on YouTube Full article