Top 3 new model launches at Gemini Audio at Night
Google DeepMind introduced three major audio model advancements: Gemini 3.8 TTS for highly personalized voice generation, Gemini 3.5 Live Translate for real-time, multi-language audio/video translation, and enhanced Gemini Live models that integrate multimodal understanding, function calls, and tool calls into developer projects.
Key takeaways
-
Gemini 3.8 TTS Voice Personalization
New text-to-speech models allow for compelling voice personalization, enabling developers to guide the voice to have specific emotions, resonance, and incorporate elements like pauses.
-
Gemini 3.5 Live Translation
A new feature enabling real-time translation for both video and audio input feeds across over 100 languages.
-
Gemini Live Models Enhancements
Updated Gemini Live models allow developers to integrate multimodal understanding, live interactions, function calls, and tool calls directly into their developer projects.