GPT-4.1 Arrives in the API: 21% Better at Coding, 83% Cheaper
Curated by the Inblix editorial team
OpenAI just dropped a new family of models into the API, and the numbers demand attention. GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano are here, and they aren’t just incremental updates—they’re a direct assault on the price-performance curve that developers obsess over. The headline figure is a 54.6% score on the SWE-bench Verified coding benchmark. That’s a 21.4 percentage point jump over GPT-4o and a 26.6 point leap past the compute-heavy GPT-4.5. If you’re shipping code, this model was trained with you in mind, specifically to stop making extraneous edits and to follow diff formats reliably.
But raw intelligence is only half the story. The economics have shifted dramatically. GPT-4.1 mini doesn’t just nip at the heels of the old flagship; it beats GPT-4o on many intelligence evals while slashing latency by nearly half and cost by 83%. For tasks where speed is the only currency that matters, the new GPT-4.1 nano is a fascinating flex—boasting a 1 million token context window in a tiny package that scores higher than GPT-4o mini on coding benchmarks like Aider polyglot. It’s a clear signal that ‘small’ no longer means ‘dumb.’
These gains translate directly to agentic workloads. OpenAI explicitly positions the 4.1 family as the backbone for agents that need to resolve customer requests or extract insights from massive documents with minimal hand-holding. The model’s ability to actually use that massive context window—setting a new state-of-the-art on the Video-MME benchmark for long-context understanding—makes it a serious tool for production environments, not just a parlor trick. Alpha testers like Windsurf, Qodo, and Thomson Reuters are already validating this in the trenches on domain-specific tasks.
There is a bit of an elephant in the room, though. GPT-4.1 is strictly an API play; you won’t find a toggle for it in ChatGPT. The consumer product will just keep absorbing these improvements silently. Meanwhile, the clock is ticking for GPT-4.5 Preview users. OpenAI is shutting it down on July 14, 2025, having learned what they needed from the ‘research preview.’ They promised to carry forward its nuanced writing and humor, but for developers paying the bills, pure performance per dollar is the punchline that matters most.
💡 Key Takeaways
- GPT-4.1 delivers a massive 21.4 percentage point absolute improvement on SWE-bench Verified over GPT-4o, specifically trained to reduce extraneous edits and follow diff formats.
- GPT-4.1 mini is not just a budget option; it actually outperforms the original GPT-4o on intelligence benchmarks while costing 83% less and cutting latency nearly in half.
- The new nano model packs a 1 million token context window into an extremely fast and cheap package, scoring 9.8% on Aider polyglot coding—beating GPT-4o mini.
- OpenAI is killing off the GPT-4.5 Preview API in three months (July 14, 2025), declaring GPT-4.1 offers similar or better performance at a fraction of the cost and speed.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.