Google DeepMind Ships Three Physical AI Models for Whole-Body Control, Dexterity, and Multi-Robot Collaboration

dexterityembodied reasoning vlmgemini robotics 2google deepmindmulti-robot collaborationphysical aivision-language-action (vla) modelwhole-body control

Google DeepMind Ships Three Physical AI Models for Whole-Body Control, Dexterity, and Multi-Robot Collaboration


Google DeepMind has released Gemini Robotics 2, the intelligence layer for its next generation of robots. This release advances beyond table-top manipulation to enable whole-body control, five-finger dexterity, and multi-robot teamwork. The system includes three distinct models, each offered under different access tiers.


Most robots today are either pre-programmed or tele-operated for narrow, repetitive task sequences. They struggle to adapt to unpredictable environments, and skills rarely transfer between different robot bodies. Gemini Robotics 2 targets all three limitations simultaneously.


TL;DR

  • Three models ship together: a Vision-Language-Action (VLA) model, an embodied reasoning Vision-Language Model (VLM), and an on-device VLA.
  • One checkpoint drives Apollo 2 with two different hand types, plus a Franka Duo gripper.
  • Gemini Robotics ER 2 is available in public preview; the VLA and on-device models remain gated.
  • Multi-finger dexterity remains the weakest area, with performance ranging from 32% to 92% success rates.
  • ASIMOV-Agentic, a new safety benchmark, is available on Hugging Face under CC-BY-4.0 licensing.

3 Models and What They Do


Gemini Robotics 2 VLA

Gemini Robotics 2 is the Vision-Language-Action (VLA) model. It converts visual and language input into direct motor control. The model can drive full humanoids from feet to fingertips, as well as other bi-arm robots. It also handles dexterous manipulation using both multi-finger hands and parallel grippers.

via MarkTechPost

Related