AI Pulse by Inblix

Google's Gemini 2 bots can now track their own screw-ups with 60% accuracy

Ars Technica AI · Jul 30, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google's Gemini 2 bots can now track their own screw-ups with 60% accuracy

Google DeepMind is pushing its robot ambitions forward with Gemini Robotics 2, a new suite of models that gives physical machines a sharper sense of what they’re doing — and what’s going wrong. The headliner is Gemini Robotics ER 2, an upgraded vision-language model that doesn’t just parse instructions but watches the robot’s own progress in real time through live camera feeds. It’s an “embodied reasoning” play that DeepMind claims is a significant leap over the 1.6 release, and developers can kick the tires starting today through an integration with the Gemini Live API.

The most concrete number Google is throwing out is a 60% accuracy rate on classifying video frame completeness. That means the model can tell if a task is actually finished or if the robot is stuck mid-step. It’s not exactly bulletproof — failing four out of ten checks isn’t something you’d want on a factory floor — but it’s a massive improvement over what came before. For context, DeepMind says both the older 1.6 model and competing visual AI systems perform far worse on this benchmark. It’s the difference between a robot that blindly executes commands and one that at least has a fighting chance of noticing it dropped the thing you asked it to pick up.

Behind the scenes, Google is chasing what its scientists sometimes call “physical AGI” — a generalist robot that can handle whatever you throw at it, not just the narrow, pre-programmed stunts we’ve seen in viral videos for years. The 2.0 release apparently controls entire humanoid forms with improved dexterity, even the kind of complex hands that make most roboticists sweat. The trio of sub-models also adds the ability to collaborate with other robots and continuously adapt to shifting environments, which moves the needle from “cool demo” to something that might actually hold down a job.

Still, it’s worth keeping the hype in check. A 60% success rate on step-completion tracking, while a genuine technical gain, underscores just how far we are from robots that can reliably operate in messy, real-world settings without human babysitters. Google’s making the ER 2 model available now, which means the developer community gets to decide whether this is a foundation for something practical or just a better set of training wheels.

💡 Key Takeaways

  1. Gemini Robotics ER 2 processes live video to track task progress, hitting 60% accuracy on frame completeness — a big relative jump but still unreliable for unsupervised work.
  2. The 2.0 release targets 'physical AGI' by controlling full humanoid robots with complex hands, not just the narrow, pre-programmed stunts of earlier machines.
  3. Developers get immediate API access to ER 2 via Gemini Live, shifting the burden of real-world testing and application-building to the open community.
  4. Multi-robot collaboration and environmental adaptation are now baked in, suggesting Google sees the path to useful robots running through teamwork, not just smarter individuals.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles