# Introducing Gemini Robotics 2

## Executive summary

Google DeepMind introduced Gemini Robotics 2, a new suite of models designed to provide the intelligence layer for general-purpose robotics. The system enables whole-body understanding and reasoning, allowing robots to perform complex tasks like cleaning a garage or folding laundry based on natural language prompts. Key advancements include enhanced dexterity, multi-robot collaboration capabilities, and leveraging Gemini's multimodal world understanding by adding 'actions' as a modality.

## Key takeaways

- Whole-Body Intelligence: Gemini Robotics 2 enables models to understand the entire robot's position in space and reason about complex, multi-step tasks (e.g., cleaning a garage), moving beyond simple object manipulation.
- Enhanced Dexterity: The models significantly improve dexterity, allowing robots to perform intricate daily tasks such as folding laundry or precisely unscrewing objects using high-DOF hands.
- Multi-Robot Collaboration: A new capability allows the robot intelligence to understand when and how to call other robots to accelerate tasks or perform actions in parallel.
- Availability and Deployment: The Embodied Reasoning (ER) model will be available via AI Studio and the Gemini Enterprise Agents Platform. An on-device version is also available through a trusted tester program.

## Technical details

- Foundation Model Architecture: The system builds upon the concept of Vision Language Action (VLA) models, enabling robots to understand natural language and visual input for general control. Gemini Robotics 2 integrates actions as a core modality into Gemini's multimodal understanding.
- Data Scaling Challenges: The primary bottleneck for physical AGI is the lack of a scalable 'internet of physical interaction data.' Data collection methods range from costly teleoperation to less labeled egocentric human video, creating an 'embodiment gap' that must be crossed.
- Safety and Benchmarking: The ER model is noted for its improved video understanding and safety features. A new open-source benchmark, 'Asimov Agentic,' has been introduced to test semantic physical understanding and common sense.
- Hardware vs. Software: The speaker notes that while locomotion is nearing a solved problem (using techniques like reinforced learning and sim-to-real transfer), dexterous manipulation remains the hardest, most contact-rich, and least generalized aspect of robotics.

## Practical implications

- The focus on general-purpose tasks (rather than narrow industrial policies) suggests that deployment will initially target semi-structured environments like industrial settings before moving to human-centric spaces.
- The development of the ER model's video understanding and task semantics means robots can now understand *how far along* a process is, enabling better agentic decision-making.
- The availability via AI Studio and an on-device version allows external partners to fine-tune models directly for specific tasks or hardware.

## Topics

General Purpose Robotics, Embodied AI, Multimodal Foundation Models (VLA), Dexterous Manipulation, Multi-Robot Systems, Reinforcement Learning, Gemini Robotics 2, ER model (Embodied Reasoning), Google AI / Gemini

Source: https://www.youtube.com/watch?v=-rYFDefcq3k
