OpenAI: AI Alignment Won't Work Without Social Scientists
Curated by the Inblix editorial team
OpenAI is making an unusually candid admission: its machine learning researchers are out of their depth when it comes to understanding the messy reality of human values. In a new paper, the lab argues that its long-term safety work—specifically algorithms designed to align AI with what people actually want—has a glaring blind spot. The problem? Humans are unreliable narrators of their own desires, riddled with cognitive biases, inconsistent ethics, and shifting preferences that break ML models trained on their feedback. The solution isn’t more powerful algorithms. It’s social scientists.
The core issue is that asking people what they want is the wrong question—or at least an incomplete one. Research has repeatedly shown that human judgment warps under pressure. People make contradictory choices about gambles when tasks get complex. Slip the word “morally” into a question, and suddenly judgments about how wrong an action is start to shift. These aren’t edge cases; they’re fundamental instabilities in the data that alignment algorithms are supposed to learn from. OpenAI’s existing techniques, including iterative amplification and AI-driven debate, were built to target the reasoning behind human values. But the lab now acknowledges they have no idea how these methods behave when real people start debating genuinely fraught, value-laden topics in natural language.
So OpenAI is proposing a workaround that’s almost retro in its simplicity: strip out the AI entirely. Instead of trying to model human debate, they want to study it directly. The concept involves running experiments where humans play the role of ML agents—two human debaters and a human judge—to map out how people actually reason through complex values. The goal is to understand the dynamics well enough that lessons from the human-only version can inform actual machine learning. This requires careful experimental design rooted in decades of cognitive science, behavioral economics, and moral psychology. As the paper puts it, most AI safety researchers are laser-focused on ML, which is “not sufficient background to carry out these experiments.”
The move signals a pragmatic shift for a lab built on scaling laws and compute. OpenAI isn’t just calling for interdisciplinary collaboration; it’s actively hiring full-time social scientists to lead this research in-house. The paper itself emerged from an ongoing workshop at Stanford’s Center for Advanced Study in the Behavioral Sciences, co-organized with scholars including Mariano-Florentino Cuéllar and Margaret Levi. It’s a recognition that the alignment problem is fundamentally a human problem—and that no amount of engineering prowess will solve it if you don’t understand the humans you’re trying to align with.
💡 Key Takeaways
- OpenAI admits its alignment algorithms may fail because human feedback is riddled with biases that shift depending on how questions are phrased, making the underlying data unreliable.
- The lab is proposing human-only experiments—where people play the role of AI agents in debate and amplification setups—to study reasoning dynamics before any machine learning is involved.
- OpenAI is actively hiring full-time social scientists, acknowledging that machine learning expertise alone is insufficient for designing experiments on human cognition and values.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.