Product
RLHF's Dirty Secret: Bigger Reward Models Don't Prevent Overoptimization
OpenAI Blog · Jul 18, 2026 · 2 min read
Everyone doing RLHF knows the score. You train a reward model on human preferences, then you optimize the hell out of y...
1 article
Explore our coverage of overoptimization — 1 curated articles, summaries, and related resources from the Inblix archive.