AI hiring tools don't just learn bias—they invent new ones
Curated by the Inblix editorial team
You might assume that the biases lurking inside AI hiring tools are simply a reflection of our own, scraped up from a flawed internet. That assumption is dangerously incomplete. New research from Princeton University and the University of Chicago shows that large language models can spontaneously cook up their own stereotypes based on a few bad experiences, and they do it far more aggressively than people.
Researchers put models including ChatGPT, Claude, and Gemini through a simulated hiring game where they had to fill 20 different jobs with candidates from four fictional ethnic groups. The twist? Every candidate was equally likely to succeed at any job. The models didn’t know that. Presented with early random failures, the LLMs quickly started segregating entire groups into specific, often lower-status, roles. When one model saw an Aima candidate fail as a doctor, it stopped hiring any Aimas for that profession and shunted them toward janitorial positions instead.
The numbers are stark. On a segregation scale where 2 represents total confinement of groups to distinct job niches, the human participants from the original psychology study scored 0.84. The AI models scored roughly 65% higher. OpenAI’s o3 reasoning model hit 1.83, nearly the ceiling. “LLMs really are eager to create generalizations from limited data,” says Ryan Liu, a Princeton PhD student and coauthor of the study. That eagerness is a feature, not a bug; these models are optimized to find patterns in math and logic puzzles, but that same instinct turns toxic in a social setting.
The problem gets thornier when you consider the industry’s push toward agentic models with persistent memory. Angelina Wang, a computer scientist at Cornell, warns that as chatbots remember past conversations to personalize your experience, they can “over-index on the same kinds of behaviors” and cement biases. Simply asking the model to be fair didn’t do much. The real fix was structural: promising the model a bonus for diverse hiring slashed bias significantly, and feeding it detailed, individual-level data like a person’s education made it less likely to lazily judge them by ethnicity. We know how to curb the worst impulses of these systems, but only if we’re willing to explicitly design for it rather than hoping a chatbot’s sense of justice kicks in.
💡 Key Takeaways
- AI models develop their own stereotypes from limited experience, scoring 65% higher on a segregation scale than humans who performed the same hiring task.
- Newer reasoning models like OpenAI’s o3 show even stronger bias because their core strength—generalizing from minimal data—backfires in social contexts.
- Simply telling an LLM to be fair is ineffective, but providing a direct incentive for diversity or feeding it rich personal data about individuals significantly reduces stereotyping.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.