Moving beyond static commands, our latest embodied reasoning model unlocks fluid task orchestration, real-time temporal intelligence, and multi-robot collaboration for the physical world.
- A “Brain” for Physical Agents: Gemini Robotics ER 2 acts as a high-level cognitive engine, seamlessly orchestrating complex, multi-step tasks by delegating physical execution to lower-level models without jarring pauses.
- Mastering Time and Space: Unprecedented temporal and spatial intelligence allows robots to track continuous task progress via live video, pinpoint exact moments to transition actions, and accurately read a wide array of physical instruments.
- Teamwork and Uncompromising Safety: Built-in safety protocols and multi-robot collaboration capabilities empower diverse machines to communicate, hand off tasks, and work securely in shared human environments.
For decades, the dream of robotics has been to create machines capable of seamlessly assisting humans in our everyday, unpredictable environments. However, the physical world is messy. For a robot to truly be helpful, accurate spatial reasoning is only half the battle. Robots must also possess the ability to think quickly, timing their decisions and reasoning with the real-time speed of life.
Today, we are taking a massive leap forward in that journey by introducing Gemini Robotics ER 2. Built on the baseline of Gemini 3.6 Flash, this is our most capable “embodied reasoning” model to date. It represents a fundamental step change in powering robots with advanced video understanding, fluid task orchestration, and the ability to collaborate natively.
Think of Gemini Robotics ER 2 as the high-level brain for your robotic fleet. It allows robots to chat naturally with humans, understand the nuances of the physical world, and plan intricate, multi-step tasks. While ER 2 handles the cognitive heavy lifting, it smoothly hands off motor execution to any lower-level vision-language-action (VLA) model. This design allows the robot to “think” about its next move while simultaneously performing its current actions, effectively bridging the gap between digital reasoning and physical execution.

Fluid Orchestration Without the Pause
Most tasks in the real world require a sequence of complex steps. Gemini Robotics ER 2 operates as a physical agent, orchestrating these steps, self-correcting when things go wrong, and generalizing its knowledge to novel situations. Developers can declare low-level control interfaces—like VLA models or navigation APIs—as tools, and stream multimodal video, audio, or text directly into the model. ER 2 can even natively call tools like Google Search to retrieve missing information or execute user-defined functions on the fly.
In the past, high-level reasoning in robotics often came at the cost of execution speed, resulting in jarring “stop-and-think” pauses. Gemini Robotics ER 2 solves this by integrating directly into the Gemini Live API. By utilizing a bidirectional streaming endpoint optimized for latency-sensitive tasks, ER 2 commands action models with fluid, uninterrupted orchestration. To demonstrate this capability in the real world, we partnered with Boston Dynamics to integrate ER 2 with Spot. By orchestrating Spot’s navigation and manipulator APIs, ER 2 transforms the quadruped into an interactive, highly responsive agent capable of fetching specific objects on command.
Unlocking Temporal and Spatial Intelligence
One of the hardest challenges in robotics is giving a machine the intuition to know exactly when a task is finished. Representing a massive upgrade over its predecessor, Gemini Robotics ER 1.6, this new model introduces foundational capabilities for task progress understanding: Continuous Progress Classification and Precision Moment-Finding.
By watching continuous raw video feeds, ER 2 tracks its own progress. Through Continuous Progress Classification, it assigns each frame of a video feed into granular completion levels (0-20%, 20-40%, etc.). This real-time situational awareness allows the robot to adjust its actions on the fly or retry a failed step without having to restart the entire workflow. Meanwhile, Precision Moment-Finding empowers the model to identify the exact video frame where a critical event takes place—such as the exact millisecond to stop pouring coffee into a cup—ensuring tasks like tightening a light bulb or tying a trash bag are completed to exact specifications.

Beyond temporal intelligence, Gemini Robotics ER 2 significantly advances our core spatial reasoning capabilities across several key benchmarks:
- Success/Failure Detection: The model now analyzes raw video feeds instead of static snapshots, allowing it to catch mid-execution failures like slips, spills, or misalignments as they happen.
- General Instrument Reading: ER 2’s comprehension extends far beyond simple circular dials. It can now accurately read digital displays, linear scales, rulers, and liquid thermometers across ten different types of physical instruments.
- Enhanced Spatial VQA: Leveraging Gemini’s broader advancements, the model boasts vastly improved Visual Question Answering for complex spatial environments.
The Power of Multi-Robot Collaboration
No single robot is perfectly suited for every task. A wheeled rover might be the most efficient choice for navigating flat indoor hallways, while a bipedal humanoid robot is better equipped for uneven terrain or human-scale manipulation.
Gemini Robotics ER 2 introduces robust multi-robot collaboration, allowing diverse machines to communicate through a shared semantic understanding. This enables different types of robots to coordinate, hand off items, and complete complex workflows that a single robot simply could not achieve alone. For example, developers can now seamlessly orchestrate workflows where Apptronik’s Apollo 2 and Franka’s F3 Duo collaborate in real-time, sharing workspace and task loads efficiently.

Advancing Safety for Embodied Intelligence
Bringing AI into the physical world requires an uncompromising commitment to safety. Gemini Robotics ER 2 is our safest model yet, achieving significant gains on Safety Instruction Following and Human Proximity benchmarks. These benchmarks evaluate how well a model adheres to physical constraints and maintains spatial awareness around people. In practical application, Gemini Robotics ER 2 will successfully halt a humanoid robot the moment a person steps too close, autonomously resuming its work only once the area is clear and safe.
To continue pushing the boundaries of safe physical agents, we are also introducing a new benchmark designed specifically to evaluate a foundation model’s ability to act as a safe VLA orchestrator. This tests the model’s capacity to strictly enforce safety constraints, monitor dynamic environments, assess the physical feasibility of a request, and proactively seek human clarification when a situation is ambiguous.
Gemini Robotics ER 2 is making it possible for robots to be genuinely helpful, adaptable, and safe in the real world. The model is now publicly available to developers via the Gemini API and Google AI Studio, and is currently in private preview on the Gemini Enterprise Agent Platform. We are incredibly excited to see what developers and roboticists will build next with this new foundation for physical AI.
