OpenAI says aligning superhuman AI now is just table stakes
Curated by the Inblix editorial team
OpenAI just laid out its alignment strategy with a refreshingly blunt admission: the techniques that work today won’t cut it for AGI. The research team, which published its approach in a detailed post, is treating alignment not as a one-time fix but as an iterative, empirical grind. The core bet? Build AI systems that are aligned enough right now to help solve the harder alignment problems that emerge later. It’s a recursive strategy — using today’s tools to build tomorrow’s safeguards.
The immediate workhorse remains reinforcement learning from human feedback, or RLHF. It’s the secret sauce behind InstructGPT, a fine-tuned version of GPT-3 that humans consistently prefer over a model 100 times its size. The economics are striking: the fine-tuning costs less than 2% of GPT-3’s original training compute and required around 20,000 hours of human feedback. That’s a massive leverage ratio, and it’s already paying off in the real world. OpenAI’s API customers are voting with their queries, overwhelmingly choosing InstructGPT over the raw pretrained models.
But the team isn’t popping champagne. They’re the first to point out the cracks. Today’s InstructGPT still fumbles simple instructions, isn’t reliably truthful, and sometimes serves up biased or toxic responses. A more subtle failure caught them off guard: some users find the aligned models significantly less creative, a problem their public benchmarks completely missed. These are the messy realities you only discover when paying customers start poking at your product in unpredictable ways.
Looking ahead, RLHF has a built-in expiration date. It requires humans to accurately judge AI outputs — a comfortable assumption when models are dumber than we are, but a shaky one when they start critiquing massive codebases or dissecting scientific papers at superhuman speed. To stay ahead of that curve, OpenAI is pouring effort into two additional pillars: training AI to assist human evaluation and, most ambitiously, training AI to do alignment research itself. The endgame is an AI that can help solve the alignment problem for all subsequent, more powerful AIs. The company says it will share this research openly when it’s safe, pushing every AGI developer to adopt the best techniques available.
💡 Key Takeaways
- InstructGPT's fine-tuning cost less than 2% of GPT-3's original training compute but produced a model that humans prefer over one 100 times larger.
- OpenAI discovered a real-world gap in its testing: paying customers find the aligned models less creative, a flaw that public benchmarks never flagged.
- The company's entire alignment strategy is recursive — it's betting that sufficiently aligned AI today can accelerate the research needed to align superhuman AI tomorrow.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.