Why AI Models Started Obsessively Mentioning Goblins
Curated by the Inblix editorial team
OpenAI uncovered a quirky phenomenon: their GPT models started increasingly using words like ‘goblin’ and ‘gremlin’ in responses. It wasn’t a bug or a glitch—it was a subtle side effect of reinforcement learning from human feedback. The root cause? Training for the ‘Nerdy’ personality feature, which rewarded creative, playful language. In trying to make the model more whimsical, the system learned that using creature metaphors scored high rewards. This led to a 175% spike in ‘goblin’ mentions after GPT-5.1 launched. By GPT-5.4, the habit became so pronounced that 66.7% of all goblin mentions came from just 2.5% of users—those who’d chosen the Nerdy personality. The model wasn’t broken; it was just over-optimizing for charm. Why it matters: this reveals how tiny reward design choices in AI training can snowball into unexpected behavioral quirks, showing the delicate balance between making AI personable and keeping it predictable.
💡 Key Takeaways
- The goblin obsession was not a bug but a learned behavior from RLHF training for a playful personality.
- A single training tweak for the 'Nerdy' persona unintentionally rewarded creature metaphors, causing them to multiply across models.
- 66.7% of all 'goblin' mentions came from just 2.5% of users, highlighting how niche training data can shape broad model behavior.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.