GPT-4 Passes the Bar, But the Real Story Is Predictability
Curated by the Inblix editorial team
OpenAI just dropped GPT-4, and if you’re looking for a headline number, here it is: the model scored in the 90th percentile on a simulated bar exam. That’s a staggering leap from GPT-3.5, which limped in at the bottom 10%. But fixating on exam scores misses what’s actually novel here. For the first time, OpenAI says it accurately predicted the training performance of a large model before the run even finished. That might sound like an internal brag, but it’s a genuinely big deal for safety research. If you can reliably forecast what a model will be capable of before you finish building it, you can start designing guardrails well in advance rather than scrambling to bolt them on after the fact. The company rebuilt its entire deep learning stack and, alongside Azure, co-designed a supercomputer specifically for this workload. The result was an “unprecedentedly stable” training run—their words, and for good reason.
The model itself is multimodal, accepting both image and text inputs, though images remain a research preview for now. That means you can show GPT-4 a photograph, a diagram, or a screenshot, and it will reason about it in context. The text capability is live in ChatGPT and the API, subject to a waitlist. OpenAI is also open-sourcing its evaluation framework, called Evals, to crowdsource the detection of model shortcomings. On traditional benchmarks, GPT-4 doesn’t just edge out the competition—it “considerably outperforms existing large language models,” including systems that were explicitly fine-tuned for those tests. The multilingual performance is particularly striking. The team translated the MMLU benchmark into 26 languages using Azure Translate, and in 24 of them, GPT-4 beat the English-language performance of GPT-3.5. That includes low-resource languages like Latvian, Welsh, and Swahili, which typically get left behind in these releases.
For casual users, the jump from 3.5 to 4 might feel subtle. The difference sharpens when tasks get complex: longer reasoning chains, nuanced instructions, creative constraints. OpenAI has been dogfooding the model internally across support, sales, and content moderation, and they frame this as the second phase of their alignment strategy—using AI to help humans evaluate AI outputs. It’s a pragmatic move, though it also raises the recursive question of who watches the watchers. The six months of adversarial testing and alignment work, informed by the ChatGPT deployment, have produced what they call their “best-ever results” on factuality and steerability. They’re careful to note it’s still far from perfect. The safety framing isn’t just boilerplate, either. The blog post explicitly ties the ability to predict training performance to a longer-term goal of anticipating future capabilities “increasingly far in advance.”
We’ve seen plenty of models that ace benchmarks. Few have been built with this level of upfront infrastructure investment and a stated obsession with predictability. The real test, as always, will be in the messy, adversarial reality of millions of users probing the edges. Image input remains gated behind a single partner collaboration, which suggests either extreme caution or a strategic rollout designed to build anticipation. Given the competitive pressure from Google and the open-source community, my money’s on caution. GPT-4 doesn’t need to be a revolution to be significant. A model that’s more reliable, more multilingual, and whose creators are finally getting serious about forecasting its behavior before it ships? That’s a meaningful shift, even if it lacks the whiz-bang theatrics of a demo.
💡 Key Takeaways
- GPT-4's bar exam score jumped from the 10th to the 90th percentile compared to GPT-3.5, but the team's ability to predict training performance before the run is the more important safety milestone.
- The model is multimodal and can reason over images, though this capability is being rolled out cautiously through a single partner rather than made publicly available.
- Multilingual performance is a standout: GPT-4 beat GPT-3.5's English-language scores in 24 of 26 languages tested, including low-resource ones like Swahili and Latvian.
- OpenAI is open-sourcing its Evals framework so anyone can report model shortcomings, while simultaneously using GPT-4 internally to help humans evaluate other AI outputs.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.