Top 3 new model launches at Gemini Audio at Night
Summary
Google DeepMind introduced three major audio model advancements: Gemini 3.8 TTS for highly personalized voice generation, Gemini 3.5 Live Translate for real-time, multi-language audio/video translation, and enhanced Gemini Live models that integrate multimodal understanding, function calls, and tool calls into developer projects.
Key takeaways
-
Gemini 3.8 TTS Voice Personalization
New text-to-speech models allow for compelling voice personalization, enabling developers to guide the voice to have specific emotions, resonance, and incorporate elements like pauses.
-
Gemini 3.5 Live Translation
A new feature enabling real-time translation for both video and audio input feeds across over 100 languages.
-
Gemini Live Models Enhancements
Updated Gemini Live models allow developers to integrate multimodal understanding, live interactions, function calls, and tool calls directly into their developer projects.
Technical details
-
Text-to-Speech (TTS)
0s
Gemini 3.8 TTS models support voice personalization, allowing control over specific emotions, resonance, and pacing (e.g., incorporating pauses).
-
Real-Time Translation
0s
Gemini 3.5 Live Translate processes both video and audio input feeds, translating content into over 100 languages.
-
Multimodal Development
0s
Gemini Live models enhance developer capabilities by allowing the integration of multimodal understanding, live interactions, function calls, and tool calls.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.