AI Pulse by Inblix

Dyna-2 trained on 170 years of human video to teach robots manipulation

MarkTechPost · Aug 13, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Dyna-2 trained on 170 years of human video to teach robots manipulation

Dyna Robotics just dropped a result that could change how robot learning works. Dyna-2, their new world-action model, was pre-trained on over one million hours of egocentric human video — roughly 170 years of continuous waking experience. That matters because robot learning has been starved for action-labelled data, which requires expensive, deliberate teleoperation. The bet: ordinary human video can substitute. So far, the bet is paying off.

The team built a data ladder with rungs at 1,000, 10,000, 100,000, and 1,000,000 hours, keeping identical proportions from each source so differences couldn’t be blamed on distribution shift. A scaling law emerged on human data — held-out MSE fit a power law with R²=0.919, and [email protected] climbed 51% across the ladder. What’s more surprising is the transfer: the same checkpoints, scored zero-shot on 39 robot tasks across two bimanual platforms, showed zero-shot action MSE dropping from 0.306·D^-0.0713 with an R² of 0.884. There’s an inflection point between 10k and 100k hours. Dyna Robotics reports that joint denoising — predicting video and action together — beat action-only on all 39 tasks at every action scale. Video prediction, it turns out, is not just a side effect. It’s the mechanism driving cross-embodiment generalization.

On real robots, the results get concrete. Each rung was post-trained on 14 tasks with at most 10 hours of robot data each, across three embodiments including 20-DOF dexterous hands. Mean normalized score rose from 20% at 1k hours to 53% at 1M hours. Lockbox Key Turning is the standout: 0% success up to 100,000 hours, then 90% at one million. Bottle Cap Untwisting, trained on roughly 10 minutes of demonstrations, still hit 50%. Against Dyna-1 — their production VLA built on Qwen3-VL-4B — an early Dyna-2 hit 1.55× the success rate pooled over 7 tasks. At unseen customer sites, Dyna-2 passed production criteria 87% of the time versus Dyna-1’s 46%.

There’s a catch, though. Dyna-2 isn’t downloadable. No public checkpoint, no API, no license. Deployment means buying a Dyna robot cell — the company’s Dyna-1 robots already run in hotels, restaurants, and laundromats. That limits the audience to mid-market service operators and multi-site enterprises with repetitive manipulation work. For research labs wanting local inference, this remains a paper, not a product. The distillation pipeline is intriguing — video sampling drops from 10,203 ms to 110 ms on one H — but without weights, the community can’t verify or build on it. Dyna Robotics seems to be betting that the moat is in the data and the system, not the open model.

💡 Key Takeaways

  1. Dyna-2's scaling law on human video transfers zero-shot to unseen robot data, with an inflection point between 10k and 100k hours of pre-training.
  2. Joint video-and-action denoising beat action-only training on all 39 robot tasks, suggesting video prediction is the key mechanism for cross-embodiment generalization.
  3. Lockbox Key Turning jumped from 0% to 90% success when pre-training scaled from 100k to 1M hours, even with only 10 hours of robot post-training data.
  4. Dyna-2 is vendor-locked: no public weights, API, or license exists, so deployment requires purchasing Dyna robot cells rather than self-hosting the model.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles