Google DeepMind Ships Three Physical AI Models for Whole-Body Control, Dexterity, and Multi-Robot Collaboration
Google DeepMind has released Gemini Robotics 2, the intelligence layer for its next generation of robots. This release advances beyond table-top manipulation to enable whole-body control, five-finger dexterity, and multi-robot teamwork. The system includes three distinct models, each offered under different access tiers.
Most robots today are either pre-programmed or tele-operated for narrow, repetitive task sequences. They struggle to adapt to unpredictable environments, and skills rarely transfer between different robot bodies. Gemini Robotics 2 targets all three limitations simultaneously.
TL;DR
- Three models ship together: a Vision-Language-Action (VLA) model, an embodied reasoning Vision-Language Model (VLM), and an on-device VLA.
- One checkpoint drives Apollo 2 with two different hand types, plus a Franka Duo gripper.
- Gemini Robotics ER 2 is available in public preview; the VLA and on-device models remain gated.
- Multi-finger dexterity remains the weakest area, with performance ranging from 32% to 92% success rates.
- ASIMOV-Agentic, a new safety benchmark, is available on Hugging Face under CC-BY-4.0 licensing.
3 Models and What They Do
Gemini Robotics 2 VLA
Gemini Robotics 2 is the Vision-Language-Action (VLA) model. It converts visual and language input into direct motor control. The model can drive full humanoids from feet to fingertips, as well as other bi-arm robots. It also handles dexterous manipulation using both multi-finger hands and parallel grippers.
via MarkTechPost
