Google Developers

Top 3 new model launches at Gemini Audio at Night

Published 2026-10-06 · Duration 1:22

Summary

Google DeepMind introduced three major audio model advancements: Gemini 3.8 TTS for highly personalized voice generation, Gemini 3.5 Live Translate for real-time, multi-language audio/video translation, and enhanced Gemini Live models that integrate multimodal understanding, function calls, and tool calls into developer projects.

Download summary

Key takeaways

  1. Gemini 3.8 TTS Voice Personalization

    New text-to-speech models allow for compelling voice personalization, enabling developers to guide the voice to have specific emotions, resonance, and incorporate elements like pauses.

  2. Gemini 3.5 Live Translation

    A new feature enabling real-time translation for both video and audio input feeds across over 100 languages.

  3. Gemini Live Models Enhancements

    Updated Gemini Live models allow developers to integrate multimodal understanding, live interactions, function calls, and tool calls directly into their developer projects.

Technical details

  • Text-to-Speech (TTS) 0s

    Gemini 3.8 TTS models support voice personalization, allowing control over specific emotions, resonance, and pacing (e.g., incorporating pauses).

  • Real-Time Translation 0s

    Gemini 3.5 Live Translate processes both video and audio input feeds, translating content into over 100 languages.

  • Multimodal Development 0s

    Gemini Live models enhance developer capabilities by allowing the integration of multimodal understanding, live interactions, function calls, and tool calls.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.