The Extended Brief

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Brief by The AI News AI newsroom · Jul 30, 2026, 5:52 PM EDT edition

Original reporting by Google DeepMind · published Jul 30, 2026, 11:00 AM EDT

DeepMind's new robotics foundation model introduces multi-robot collaboration and advanced video understanding, setting a new baseline for embodied AI systems.

Key points

  • Gemini Robotics ER 2 enables autonomous reasoning and real-world task execution for robots.
  • The system delivers significant advancements in robotic video understanding capabilities.
  • It facilitates advanced tool orchestration and multi-robot collaboration for complex applications.

From the source

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API , Google AI Studio , and in private preview on Gemini Enterprise Agent Platform .

By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step.

We are also introducing multi-robot collaboration, enabling robots to work together in shared spaces and complete complex workflows a single robot could not do alone.

We found that Gemini Robotics ER 2 successfully halts a humanoid robot when a person is nearby and autonomously resumes work only once the area is clear.

In our evaluations, we assign each frame in a video feed into five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%).

Quoted verbatim from the original article at Google DeepMind

Practical applications

  • Robotics teams can evaluate Gemini Robotics ER 2 as a reasoning layer for task planning and real-world execution in their stacks.
  • Researchers can test the model's video-understanding claims against their own embodied-perception benchmarks.
  • Teams running fleets of robots can prototype the multi-robot collaboration features for coordinated tasks like warehouse workflows.
  • Engineers building tool-orchestration pipelines can compare ER 2's orchestration against their current planner or VLA setup.

Who should care

Robotics researchers and engineers, and teams building embodied-AI products that need perception, task planning, or coordination across multiple robots.

Context

Robotics foundation models aim to give robots general reasoning and perception rather than task-specific programming, and Google DeepMind's Gemini Robotics line applies its Gemini models to that problem. ER 2 is the next step, emphasizing video understanding, autonomous reasoning for real-world task execution, tool orchestration, and — new for the space — collaboration among multiple robots. Multi-robot coordination has been a longstanding challenge, so a foundation model treating it as a first-class capability marks a shift in what embodied-AI baselines are expected to cover.

What to watch

  • Published benchmarks or third-party demos quantifying ER 2's gains in video understanding and multi-robot coordination over prior systems.
  • Availability details — API access, hardware partners, or real deployments — showing whether the model moves beyond research settings.

Editorial score 3.8 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 3.0

Desks: Research · Engineering · Tags: models, research

Evidence basis: Reviewed from a feed excerpt

This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.