Garp Independent AI & technology journalism
Friday, August 7, 2026 Sign In · Join Subscribe
Latest Defense tech Hadrian raises $1.37B at $8B valuation

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Robotics  /  Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Robotics

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and…

Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration — making it possible for robots to be more helpful in the physical world. We are launching Gemini Robotics ER 2, a new model designed to act as a high-level brain for robots.

It enables real-time spatial reasoning, multi-step task planning, and collaboration between different robots. You can access the model now via the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform to start building your own physical AI agents. Google just launched Gemini Robotics ER 2, which acts like a smart brain for robots. It helps them think quickly, plan multi-step tasks and even work together with other robots. The model lets robots watch video feeds to track their own progress and fix mistakes in real time. It’s designed to make robots safer and more helpful in the real world. For robots to assist humans in everyday environments, accurate spatial reasoning is not enough. Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world. That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics. Think of Gemini Robotics ER 2 as a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower level vision-language-action (VLA) model. Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function. The design of Gemini Robotics ER 2 allows the robot to “think” about what comes next while simultaneously performing its actions. Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6.