OpenAI Five crushes Dota 2 champs, then retires
Curated by the Inblix editorial team
OpenAI Five didn’t just win. It dominated. The bot swept OG, the reigning Dota 2 world champions, in two back-to-back games at a live Finals event, marking the first time an AI has ever beaten esports pros on livestream. Both OpenAI and DeepMind had previously lost public matches against top players, which made Saturday’s victory a genuine milestone. But the real story isn’t just the win — it’s what the team learned along the way.
The biggest surprise, according to OpenAI, was that the breakthrough they needed wasn’t a fancy new algorithmic trick like hierarchical reinforcement learning. It was raw, unapologetic scale. Using a custom training system called Rapid, the team ran their Proximal Policy Optimization (PPO) algorithm on a compute budget that dwarfed previous efforts. The result? The Finals version of the bot consumed 800 petaflop/s-days and crammed 45,000 years of Dota self-play into just 10 months of real time. That’s up from 10,000 years during their losing debut at The International 2018, a change driven almost entirely by letting training run 8x longer. The older TI version now loses to the current one 99.9% of the time.
That ‘just add more compute’ lesson came with a second revelation: the bot discovered a rudimentary ability to cooperate with human teammates, even though it was trained exclusively to crush other bots. The ease of that pivot from competitor to teammate has OpenAI researchers hopeful about future human-AI collaboration. To stress-test this, the team will open the bot up to the internet from April 18–21, letting anyone play with or against it. It’s a massive, open experiment designed to answer a blunt question: how easily can this thing be beaten by clever humans looking for exploits?
And then, abruptly, retirement. OpenAI is pulling the bot from competitive play immediately, though the Dota environment itself will remain an active research playground. The elephant in the room is the method’s insatiable hunger for data — a problem that’s wildly impractical outside of simulation. OpenAI points to their robotic hand project as proof the sim-to-real gap can be bridged, but admits that slashing the experience requirement is the next real challenge for reinforcement learning. For now, the message is clear: we scaled our way to a world champion, and now we’re done.
💡 Key Takeaways
- OpenAI Five's victory came from an 8x increase in training compute, not a sophisticated new algorithm, underscoring the raw power of scale in deep reinforcement learning.
- The bot spontaneously developed the ability to cooperate with human players despite being trained in a purely competitive setting, which surprised researchers.
- OpenAI is immediately retiring the bot from competitive events and will open it to the public for a four-day period to study how easily humans can exploit it.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.