GPT-2 Learns to Cheat on Summaries, Copying Text to Please Human Judges
OpenAI Blog · Jul 19, 2026 · 2 min read
Training AI with human feedback sounds like a straightforward path to better, safer models. But new research from OpenA...
11 articles
Explore our coverage of RLHF — 11 curated articles, summaries, and related resources from the Inblix archive.
OpenAI Blog · Jul 19, 2026 · 2 min read
Training AI with human feedback sounds like a straightforward path to better, safer models. But new research from OpenA...
OpenAI Blog · Jul 19, 2026 · 3 min read
OpenAI just proved that training models with human feedback can dramatically outperform simply scaling them up. Their 1...
OpenAI Blog · Jul 18, 2026 · 2 min read
It's a rare day when a model more than 100 times smaller wins a straight-up preference vote, but that's exactly what Op...
OpenAI Blog · Jul 18, 2026 · 2 min read
OpenAI just laid out its alignment strategy with a refreshingly blunt admission: the techniques that work today won't c...
OpenAI Blog · Jul 18, 2026 · 2 min read
Everyone doing RLHF knows the score. You train a reward model on human preferences, then you optimize the hell out of y...
OpenAI Blog · Jul 18, 2026 · 2 min read
OpenAI has dropped ChatGPT, a new conversational model that’s free to use during its research preview. It’s a sibling t...
OpenAI Blog · Jul 18, 2026 · 3 min read
It’s a strange kind of irony. Iceland, an island of roughly 370,000 people with a booming tech scene, is seeing its anc...
OpenAI Blog · Jul 17, 2026 · 2 min read
OpenAI’s Superalignment team just dropped a paper that flips the script on a core AI safety problem: how do you control...
OpenAI Blog · Jul 16, 2026 · 2 min read
OpenAI researchers have essentially taught GPT-4 to police itself. They've developed a new model, CriticGPT, specifical...
OpenAI Blog · Jul 16, 2026 · 2 min read
OpenAI is quietly overhauling how it teaches large language models to behave, swapping out some slow, expensive human f...
Inblix · Jul 10, 2026 · 1 min read
Reinforcement learning (RL) is the branch of machine learning where an AI agent learns by interacting with an environme...