Topic

Real-Time Translation

All digests tagged Real-Time Translation

Top 3 new model launches at Gemini Audio at Night thumbnail

· 1:22

Top 3 new model launches at Gemini Audio at Night

Google DeepMind introduced three major audio model advancements: Gemini 3.8 TTS for highly personalized voice generation, Gemini 3.5 Live Translate for real-time, multi-language audio/video translation, and enhanced Gemini Live models that integrate multimodal understanding, function calls, and tool calls into developer projects.

Key takeaways

  1. Gemini 3.8 TTS Voice Personalization

    New text-to-speech models allow for compelling voice personalization, enabling developers to guide the voice to have specific emotions, resonance, and incorporate elements like pauses.

  2. Gemini 3.5 Live Translation

    A new feature enabling real-time translation for both video and audio input feeds across over 100 languages.

  3. Gemini Live Models Enhancements

    Updated Gemini Live models allow developers to integrate multimodal understanding, live interactions, function calls, and tool calls directly into their developer projects.

Watch on YouTube Full article