AI Pulse by Inblix

OpenAI's new transcription models drop error rate to 3.31% but still trail ElevenLabs

The Decoder · Jul 29, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's new transcription models drop error rate to 3.31% but still trail ElevenLabs

OpenAI just released two new speech-to-text models, GPT Transcribe and GPT Live Transcribe, and the numbers paint a clear picture: solid progress, but no leaderboard shakeup. GPT Transcribe handles pre-recorded audio and chews through files at 34 times real-time speed. GPT Live Transcribe targets streaming use cases where latency actually matters. Both are available through the API now.

On the accuracy front, the independent benchmark AA-WER puts GPT Transcribe at a 3.31 percent word error rate. That’s a 0.7 percentage point improvement over the year-old GPT-4o Transcribe — nothing seismic, but directionally correct. More interesting is the pricing. OpenAI dropped the cost by 25 percent to $0.0045 per minute, which makes the upgrade a no-brainer for anyone already in their ecosystem. The models also accept text context, keyword boosting, and multiple languages, features that bring them closer to parity with what competitors have offered for a while.

The problem? Everyone else got better too. ElevenLabs Scribe v2 still sits at the top with a 2.3 percent error rate. Google’s Gemini 3 Pro clocks in at 2.9 percent, and Mistral’s Voxtral Small hits 3 percent — all ahead of OpenAI. And Mistral is playing the pricing game hard, undercutting the field with Voxtral Transcribe V2 starting at just $0.003 per minute. So OpenAI improved, but it’s chasing from behind on both accuracy and cost.

These models also complement OpenAI’s recently announced Realtime model generation, which includes GPT-Realtime-Whisper. It’s a sensible integration play. If you’re building voice agents or real-time applications on OpenAI’s stack, having transcription baked into the same ecosystem reduces complexity. But if raw accuracy or per-minute cost is your only concern, the market has better options right now. The question is whether tight integration wins over spec sheets.

💡 Key Takeaways

  1. OpenAI's GPT Transcribe hits a 3.31% word error rate, improving 0.7 percentage points over its predecessor but remaining behind competitors like ElevenLabs (2.3%) and Google (2.9%).
  2. The 25% price cut to $0.0045 per minute makes the upgrade obvious for existing OpenAI customers but doesn't beat Mistral's aggressive $0.003 per minute floor.
  3. Integration with OpenAI's Realtime stack gives these models an ecosystem advantage that pure accuracy benchmarks don't capture, which may matter more for developers building voice agents.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

← Back to all articles