OpenAI drops o1-preview: A PhD-level reasoner that thinks first
Curated by the Inblix editorial team
OpenAI isn’t calling it GPT-5. They’re resetting the counter to 1. The company just released o1-preview, the first in a new series of models trained to “spend more time thinking through problems before they respond, much like a person would.” If you’ve ever watched GPT-4o confidently barrel into a wrong answer on a hard logic puzzle, you’ll understand exactly why this shift matters. These models learn to refine their process, try different strategies, and catch their own mistakes before opening their mouths.
The benchmark jumps are jarring. On a qualifying exam for the International Mathematics Olympiad, GPT-4o solved 13% of problems. o1 hit 83%. That’s the difference between a struggling undergrad and a contender. In Codeforces competitions, it landed in the 89th percentile. OpenAI says the next update—already in development—performs like a PhD student on physics, chemistry, and biology benchmarks. Healthcare researchers annotating cell sequencing data, physicists generating quantum optics formulas, developers building multi-step workflows: those are the audiences OpenAI is targeting here, not the casual chat user.
There’s a safety story here that’s equally striking. Because the model can reason about safety rules in context, it applies them more effectively when someone tries to jailbreak it. On OpenAI’s hardest jailbreaking test, GPT-4o scored a dismal 22 out of 100. o1-preview scored 84. That’s not an incremental improvement—it’s a fundamentally different posture toward adversarial prompts. The company also formalized agreements with the U.S. and U.K. AI Safety Institutes, granting early access to a research version of the model to establish evaluation processes before and after public release.
But this is still a preview, and it shows. o1 doesn’t browse the web, can’t upload files or images, and lacks function calling, streaming, or system message support in the API. For many everyday tasks, GPT-4o remains the more capable tool. Rate limits are tight too: 50 queries per week for o1-preview, 50 per day for the smaller o1-mini. Developers at API tier 5 get 20 RPM. The model that thinks longer also makes you wait longer—and costs more. o1-mini is 80% cheaper and built for coding specifically, but sacrifices broad world knowledge. The real question isn’t whether this is impressive. It’s whether the average user has problems hard enough to justify the wait.
💡 Key Takeaways
- o1-preview scored 83% on an International Math Olympiad qualifier where GPT-4o managed just 13%, a leap that redefines what's possible on hard reasoning tasks.
- The model's ability to reason about safety rules in context pushed its jailbreaking resistance score from 22 (GPT-4o) to 84, a shift with real implications for deployment safety.
- OpenAI is releasing a smaller, 80% cheaper o1-mini variant optimized for coding, acknowledging that not every reasoning task requires the full model's breadth of world knowledge.
- The preview lacks core ChatGPT features like web browsing, file uploads, and API function calling, making GPT-4o still the practical choice for most common use cases today.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.