OpenAI bots now beat humans at Dota 2 after 180 years of daily practice
Curated by the Inblix editorial team
Five neural networks working as a team have started beating amateur human squads at Dota 2, and OpenAI plans to put them up against top professionals at The International this August. The system, called OpenAI Five, is a massive scaled-up version of the Proximal Policy Optimization algorithm the lab used for its 1v1 bot last year — now running across 256 GPUs and 128,000 CPU cores. Each hero gets its own LSTM network, and the whole thing learns entirely through self-play, churning through 180 years of games every single day without ever touching human replay data. No search, no bootstrapping, just raw reinforcement learning.
Dota 2 makes chess and Go look like tidy little board games. A typical match runs 45 minutes at 30 frames per second — 80,000 ticks — and OpenAI Five observes every fourth frame, leaving 20,000 moves per game. Compare that to chess, which rarely cracks 40 moves. The action space gets discretized into 170,000 possible moves per hero, and about 1,000 of those are actually valid at any given tick. On top of that, you’re playing with a fog of war. Units can only see what’s in their immediate vicinity, so the AI has to make inferences about enemy positions and strategy from incomplete data — something neither chess nor Go’s full-information boards ever demand.
The lab itself sounds almost surprised this works. OpenAI notes that the results “indicate that reinforcement learning can yield long-term planning with large but achievable scale — without fundamental advances, contrary to our own expectations upon starting the project.” That line matters. It suggests we might not need some theoretical breakthrough to get AI systems that handle messy, continuous, partially-observed environments. You just need enough compute and the patience to let them play against themselves for the equivalent of several human lifetimes.
A public match against top players is set for August 5th, with a Twitch stream and an in-person event for those who can snag an invite. The real question isn’t whether OpenAI Five can win — it’s whether the skills it picks up in Dota’s chaotic, high-dimensional sandbox actually transfer anywhere useful. That’s the bet behind the whole video game milestone approach: that systems which master something this messy will end up being general-purpose in ways a chess engine never was.
💡 Key Takeaways
- OpenAI Five trains purely through self-play with no human data, yet develops recognizable strategies using only reinforcement learning at massive scale.
- The system processes 20,000 moves per Dota game compared to under 40 for chess and under 150 for Go, operating in a partially-observed environment with 170,000 possible actions per hero.
- OpenAI admitted the results surprised them, stating the project succeeded "without fundamental advances" — suggesting brute-force compute can achieve what they thought required algorithmic breakthroughs.
- The August 5th match against top professionals at The International will test whether self-play alone can close the gap to world-class human teams.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.