Product
Soft Q-learning and policy gradients are secretly the same
OpenAI Blog · Jul 20, 2026 · 2 min read
Reinforcement learning has a weird schism at its heart. On one side, you've got policy gradient methods like A3C, which...
2 articles
Explore our coverage of policy-gradients — 2 curated articles, summaries, and related resources from the Inblix archive.
OpenAI Blog · Jul 20, 2026 · 2 min read
Reinforcement learning has a weird schism at its heart. On one side, you've got policy gradient methods like A3C, which...
OpenAI Blog · Jul 20, 2026 · 2 min read
OpenAI's latest Baselines release buries a quiet admission: the "asynchronous" part of A3C, the influential algorithm t...