Wikiwand AI

Gemini Robotics

Vision-language-action model From Wikipedia, the free encyclopedia

Gemini Robotics is an advanced vision–language–action model developed by Google DeepMind[1][2]. It is based on the Gemini 2.0 large language model.[3] It is tailored for robotics applications and can understand new situations.[4][5] There is a related version called Gemini Robotics–ER, which stands for embodied reasoning.[3] The two models were launched on March 12, 2025.[5]

On June 24, 2025, Google DeepMind released Gemini Robotics On-Device, a variant designed and optimized to run locally on robotic devices.[6]

On July 30, 2026, Google DeepMind announced Gemini Robotics ER 2, an upgrade to its embodied reasoning model featuring real-time video understanding, task progress tracking, lower-latency orchestration via the Gemini Live API, tool integration, and multi-robot collaboration. With this release, the model was made publicly available to developers via the Gemini API and Google AI Studio.[7]

Access to earlier Gemini Robotics models was initially restricted to trusted testers, including Agile Robots, Agility Robotics,[8] Boston Dynamics, and Enchanted Tools.[2]

References

Related Articles

Timelines

Top Qs

Fact Checks