# Workshop: Async, Sync, or Realtime: Which model do I choose?

## Executive summary

This webinar provides a comprehensive framework for selecting the appropriate Speech-to-Text (STT) model—Async, Sync, Realtime, or Dictation—based on a product's specific requirements for latency, accuracy, and cost. The choice hinges on whether the audio is pre-recorded (Async) or if a user is actively waiting for the text (Realtime/Sync). For advanced applications like voice agents, specialized models such as Universal 3.6 Pro are recommended for handling turn-taking context and background noise. Developers are cautioned against common pitfalls, such as using Async for live experiences or arbitrarily chunking audio, which can severely degrade transcription quality.

## Key takeaways

- API Selection Framework: The core decision point is whether a user is waiting for the text in real-time. If no one is waiting, Async is suitable. If someone is waiting, Realtime or Sync must be considered.
- Voice Agent Requirements: For voice agents (where the system talks back), use the state-of-the-art Universal 3.6 Pro model. This model handles turn-passing context, conversation memory, and background noise filtering.
- Dictation vs. Realtime: Dictation is designed for short clips and provides both the verbatim transcript and an LLM-cleaned-up version in a single request, simplifying the process compared to multiple requests.
- Avoiding Chunking Pitfalls: When processing audio, do not chunk arbitrarily (e.g., 30-second clips). Always chunk based on natural utterances (spoken phrases) to maintain high word error rate accuracy.

## Technical details

- API Workflow Comparison: Async processes large, pre-recorded files (up to 10 hours) via webhooks/pull. Sync handles short, chopped clips (max 2 minutes) with a quick request/response. Realtime streams words via a Websocket connection. Dictation is a single request providing both verbatim and cleaned text.
- Model Differentiation (Pro vs. Base): The Universal 3.6 Pro model is for complex, interactive scenarios (voice agents) requiring advanced features like turn detection and context passing. The Universal 3.6 Base model is simpler, cheaper, and optimized for pure transcription tasks like notetaking or live captions.
- Latency Benchmarks: Sync and Dictation offer near-instantaneous round-trip latency (approx. 1 second). Realtime provides initial words within 1 second of connection, with final text taking 3-4 seconds after audio ends. Async latency is typically 10 to 32 seconds for a full file.
- Advanced Features and Pitfalls: Speaker diarization (identifying who spoke what) is available as an add-on feature for Async. Pitfalls include: 1) Using Async when real-time interaction is needed. 2) Sending pre-recorded files through the Realtime stream (use Async instead). 3) Arbitrarily chunking audio instead of using a Voice Activity Detector (VAD) and sending to Sync.

## Practical implications

- When designing a product, prioritize the user experience (latency) over the lowest cost, especially for live features.
- For any voice agent or automated phone line, utilize the Pro model to ensure accurate turn-passing context and noise filtering.
- If building a real-time flow but needing to process chunks, use the Sync API with a VAD-generated chunking strategy rather than the Realtime Websocket.
- For call recordings or QA, Async is the most efficient path, offering high accuracy and supporting speaker diarization.

## Topics

Speech-to-Text (STT), Asynchronous Processing, Realtime Streaming, Voice Agent Development, Latency Optimization, Websockets, AssemblyAI Website

Source: https://www.youtube.com/watch?v=vPs3DLyfNew
