AI Pulse by Inblix

FrontierScience tests how well AI can do real science

OpenAI Blog · Jul 11, 2026 · 1 min read · Read original article →

Curated by the Inblix editorial team


OpenAI dropped a new benchmark called FrontierScience to see if AI models can actually do expert-level scientific reasoning in physics, chemistry, and biology. Unlike older benchmarks that just test multiple-choice fact recall, FrontierScience has two tracks: Olympiad-style problems and open-ended research tasks that mimic real scientific work. The results show GPT‑5.2 leading the pack with 77% on the Olympiad track but only 25% on the Research track — a sign that AI still struggles with the messy, creative side of science. Meanwhile, GPT‑5 jumped from a 39% score on the GPQA benchmark two years ago to a staggering 92% now, proving models are getting way better at hard science questions. The big deal here is that FrontierScience is built by PhD experts and designed to be tough enough that even top models have room to grow. Why it matters: This benchmark finally gives us a meaningful way to measure whether AI can actually help scientists speed up discoveries, not just ace trivia tests.

💡 Key Takeaways

  1. FrontierScience tests AI on both Olympiad-style reasoning and real-world research tasks, with current models scoring much lower on the latter.
  2. GPT‑5.2 leads frontier models but still only scores 25% on research-style questions, highlighting a gap in open-ended scientific problem-solving.
  3. AI progress in science is accelerating fast — GPT‑5.2 hit 92% on GPQA, up from GPT‑4's 39% two years earlier.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles