AI Pulse by Inblix

OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score

OpenAI Blog · Jul 15, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score

OpenAI just open-sourced a new factuality benchmark, and the results are humbling. The dataset, called SimpleQA, is designed to be brutally straightforward: 4,326 short, fact-seeking questions with single, indisputable answers. No room for debate. No points for style. And GPT-4o, one of the strongest models available, scores below 40%.

The team built SimpleQA to solve a measurement problem. Factuality is notoriously slippery to evaluate because models generate long, messy completions. By constraining the task to concise Q&A pairs, OpenAI made grading tractable. They had AI trainers browse the web to create questions, then had a second trainer answer them blind. Only pairs with agreement made the cut. A third trainer audited 1,000 questions and matched the original answers 94.4% of the time, with the team estimating a roughly 3% inherent error rate in the dataset.

What makes this benchmark genuinely useful is that it’s not saturated. Older benchmarks like TriviaQA are effectively solved. SimpleQA was explicitly designed to trip up frontier models—questions were selected in part because they induced hallucinations from GPT-4o and GPT-3.5. The diversity is real too, spanning science, tech, TV shows, and video games.

The calibration findings are equally interesting. Models in the o1 family, which are designed to spend more computation thinking, choose to say “not attempted” far more often than their GPT-4o counterparts. They appear to use reasoning to recognize gaps in their knowledge rather than bullshitting. That’s progress, but it also means these models are less useful when you need an answer, any answer. The tradeoff between precision and coverage isn’t going away anytime soon.

💡 Key Takeaways

  1. GPT-4o answers fewer than 40% of SimpleQA questions correctly, confirming that even top-tier models still struggle with basic factuality when the benchmark isn't saturated.
  2. The o1 model family deliberately opts to not answer questions rather than hallucinate, suggesting reasoning compute can be routed toward self-assessment and uncertainty estimation.
  3. SimpleQA's 3% estimated error rate comes from a rigorous triple-trainer validation process, making it one of the cleaner factuality benchmarks available for researchers.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles