OpenAI found 80% public agreement on AI rules — until politics came up
Curated by the Inblix editorial team
OpenAI just put a number on something that’s usually just vibes: when you ask regular people how an AI should behave, they agree with the company’s Model Spec about 80% of the time. The finding comes from a collective alignment experiment where over 1,000 participants worldwide ranked AI responses to synthetic prompts. Their preferences were then measured against a Model Spec Ranker — essentially GPT‑5 Thinking instructed to apply the Spec’s rules to the same scenarios.
The broad alignment held up well on principles like honesty, expressing uncertainty, and avoiding overstep. But the consensus cracked predictably along culture-war fault lines. Political content, sexual or graphic material, and critiques of pseudoscience or conspiracy theories were where participant rankings most frequently diverged from the Spec’s guidance. That’s not surprising. These are the exact domains where reasonable people disagree sharply, and where any single default behavior will inevitably frustrate someone.
OpenAI is treating this as infrastructure, not a one-off PR exercise. The team transformed crowd feedback into concrete proposals, split between clarifications — cases where the Spec’s intent was right but the wording was fuzzy — and change-of-principles, where participant preferences genuinely conflicted with the existing rules. Some changes were adopted immediately. Others got deferred or shelved based on feasibility or principle. They’ve also published the underlying public inputs dataset on HuggingFace so other researchers can pick at it. That’s a genuinely useful move, especially given how opaque most alignment work remains.
The real tension here is one the post acknowledges directly: defaults are powerful, but no single behavior set will ever satisfy everyone. Personalization and custom personalities are the long-term escape hatch. In the meantime, the company is trying to build a process — eliciting preferences, translating them into behavioral guidance, updating the Spec — that can be repeated as models grow more capable and the stakes get higher. The next Model Spec update, incorporating these changes, is promised soon.
💡 Key Takeaways
- Crowd preferences matched the Model Spec roughly 80% of the time, with strong agreement on honesty and humility but sharp divergence on political and graphic content.
- OpenAI classified disagreements into 'clarifications' where wording was the problem and 'change-of-principles' where the Spec itself needed updating, adopting some changes immediately.
- The company open-sourced its public inputs dataset on HuggingFace, giving external researchers a rare window into how collective alignment data is collected and interpreted.
- Defaults matter enormously, but OpenAI concedes no single behavior set will work for everyone — personalization is the long-term strategy for handling irreconcilable preferences.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.