AI Pulse by Inblix

DeepMind's Gemini Robotics 2 gives humanoids legs and fingers, but still fumbles 68% of knot-tying

MarkTechPost · Jul 30, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: DeepMind's Gemini Robotics 2 gives humanoids legs and fingers, but still fumbles 68% of knot-tying

Google DeepMind just shipped Gemini Robotics 2, and it’s not a single model. It’s three. The company is splitting robot intelligence into a vision-language-action (VLA) model for motor control, an embodied reasoning VLM for high-level planning, and a smaller on-device VLA that runs locally. The reasoning model, ER 2, is available in public preview. The other two remain gated.

This release is about moving past tabletop demos. For the first time, the stack drives whole-body control on Apptronik’s Apollo 2 humanoid. You can tell it to “put the watering can into the green bin on the bottom shelf,” and it will walk over, pick it up, and place it. One model checkpoint works across Apollo 2 with two different hands and a Franka Duo with a parallel gripper. That cross-embodiment flexibility is the real flex here.

But the numbers tell a different story on dexterity. The same model that seals a ziplock bag 92% of the time only ties a knot successfully 32% of the time. That’s a massive spread, and DeepMind is upfront about it: multi-finger manipulation is still the weak axis. They’re not smoothing over the gap, which is refreshing.

There’s also a new safety benchmark called ASIMOV-Agentic, dumped on Hugging Face under CC-BY-4.0. The ER 2 model can now stream bidirectionally through the Gemini Live API, which should kill the stop-and-think pauses that make robot demos feel glacially slow. It’s also getting better at knowing when to stop pouring coffee—91.3% accuracy with under a second of lag. That’s the kind of temporal awareness that usually gets ignored in favor of flashier benchmarks.

💡 Key Takeaways

  1. One model checkpoint drives three different robot embodiments, including a full humanoid and a bi-arm system, which is rare in robotics deployments.
  2. Dexterity remains brittle—knot-tying success is just 32% while ziplock bag sealing hits 92%, showing how uneven fine motor skills still are.
  3. The new ASIMOV-Agentic safety benchmark is publicly available on Hugging Face under a permissive license, a concrete step toward shared safety evaluation.
  4. ER 2’s progress classification accuracy sits at only 57.4%, meaning the model is wrong about task completion status nearly half the time.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles