Build voice-first apps with Gemini 3.5 Transcribe
Google launched Gemini 3.5 Transcribe, an advanced LLM-based model designed for building voice-first applications. This model is available via both the Interactions API and the Live API, offering fast, contextually accurate transcription of multi-speaker recordings. Key strengths include superior recognition of structured data like email addresses and phone numbers, as well as robust support for over 70 different languages.
Key takeaways
-
Model Availability
Gemini 3.5 Transcribe is available on both the Interactions API and the Live API.
-
Structured Data Recognition
The LLM-based model excels at transcribing alphanumerics, such as email addresses (e.g., thorwebdev@google.com) and recognizing correct US phone number formats.
-
Multi-Language Support
The model can recognize and transcribe over 70 different languages, even when language hints are set to English.