AI Pulse by Inblix

Bigger AI models lie more, new TruthfulQA benchmark finds

OpenAI Blog · Jul 18, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Bigger AI models lie more, new TruthfulQA benchmark finds

Here’s a finding that should give anyone deploying large language models serious pause: the bigger the model, the more likely it is to generate false answers. That’s the uncomfortable conclusion from a new benchmark called TruthfulQA, designed specifically to test whether AI systems tell the truth rather than just sounding plausible.

The benchmark consists of 817 questions across 38 categories — health, law, finance, politics — crafted around misconceptions that real humans often get wrong. Think questions where a false belief is common enough that it’s baked into the internet’s training data. The best-performing model managed to answer truthfully just 58% of the time. For context, human participants hit 94%. That gap isn’t just large. It’s a chasm.

Researchers tested GPT-3, GPT-Neo/J, GPT-2, and a T5-based model, and discovered something that flips the usual scaling narrative on its head. Across most NLP benchmarks, larger models perform better — more parameters, more accuracy. But here, the largest models were generally the least truthful. The pattern makes a grim kind of sense: if falsehoods are pervasive in the training data scraped from the web, then a model that’s better at absorbing its training distribution will also be better at reproducing those falsehoods with unnerving confidence.

The paper’s central argument is that imitation learning — training models to mimic human text from the internet — has a built-in ceiling when it comes to truthfulness. Scaling up won’t fix it, because you’re just scaling up exposure to the same patterns of misconception and misinformation the benchmark was built to expose. The authors suggest that fine-tuning with objectives beyond simple imitation is a more promising path forward, though they don’t pretend to have that problem solved yet. For anyone building products on top of these models, the implication is stark: a more eloquent AI might also be a more deceptive one.

💡 Key Takeaways

  1. Larger language models were generally less truthful than smaller ones, inverting the typical pattern where scale improves performance.
  2. The best AI model answered truthfully on only 58% of questions, compared to 94% for human participants — a 36-point gap.
  3. Models frequently generated false answers that mimic popular human misconceptions, meaning they risk deceiving users who trust fluent-sounding responses.
  4. The authors argue that scaling up imitation learning from web text is a dead end for truthfulness, and new training objectives are needed.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles