AI Pulse by Inblix

GPT‑5 arrives: Better reasoning, fewer hallucinations

OpenAI Blog · Jul 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: GPT‑5 arrives: Better reasoning, fewer hallucinations

For a model that isn’t even a year removed from GPT‑4o, the leap is audible. OpenAI pulled the curtain back on GPT‑5 today, and the first thing you notice isn’t the benchmark scores — it’s the silence. Fewer hallucinations. Less confident-sounding nonsense. The company is framing this as a reasoning-first upgrade, not just a bigger data dump, and early internal tests suggest it’s working.

In a technical blog post, OpenAI revealed that GPT‑5 scores 94.6% on GPQA Diamond, a set of graduate-level science questions that used to make models sweat. That’s a 15-point improvement over GPT‑4o. On AIME 2025, a brutal math competition, it hit 92.3%. These aren’t marginal gains. Sam Altman called it “the smartest model we’ve ever built,” and for once, the numbers back the swagger. The model was trained with a technique OpenAI calls “deliberative alignment,” which teaches it to think through safety policies step-by-step before answering — essentially baking the chain-of-thought into its ethical reasoning.

What’s genuinely new here is how the model handles ambiguity. Instead of guessing when it’s unsure, GPT‑5 is designed to recognize the limits of its own knowledge. It’s a subtle shift in personality. Less eager to please, more willing to push back or ask clarifying questions. That’s a direct response to the hallucination problem that has dogged every previous release, and according to internal red-teaming data, it cuts factual errors on adversarial prompts by nearly half.

Pricing and rollout details remain vague. The post mentions a staggered release starting with Pro users, but no date. Developers will get API access “in the coming weeks,” with higher per-token costs for the larger context windows. The real test won’t be GPQA scores — it’ll be whether GPT‑5 can maintain this level of restraint when millions of users start poking at it. Smarter models have a habit of finding creative new ways to be wrong.

💡 Key Takeaways

  1. GPT‑5 achieves a 94.6% on GPQA Diamond, a 15-point improvement over GPT‑4o that signals genuine progress in graduate-level scientific reasoning.
  2. OpenAI is using "deliberative alignment" — making the model think through safety policies before responding — to reduce harmful outputs and hallucinations.
  3. The model is designed to express uncertainty rather than fabricate answers, cutting factual errors on adversarial prompts by nearly half in internal tests.
  4. The staggered rollout starts with Pro users at an unannounced date, with higher API costs expected for developers using larger context windows.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles