ASR leaderboard scores are misleading — here's what actually matters when picking a model in 2026
Curated by the Inblix editorial team
The Open ASR Leaderboard is no longer a ranking you can trust at face value. The top models — Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR-3B, and MOSS-Transcribe-preview — are now separated by less than one word error rate point, but that number is built on shifting sand. Some averages include the easier TED-LIUM dataset and some don’t. Recompute Cohere’s published 5.42% WER over the same seven sets ARK uses, and it lands at 5.84%. Granite moves from 5.33% to 5.65%. You cannot subtract one published figure from another and get a meaningful answer.
Then there’s the benchmark gaming. The MOSS-Transcribe-preview-2B model card openly states the model was fine-tuned with reinforcement learning on the leaderboard’s own training splits. That’s disclosed, which is honest, but it means the score measures the benchmark, not real-world capability. The private-track data from Appen tells a different story entirely: toggle on held-back evaluation sets covering Australian, Canadian, Indian, and American accents in spontaneous conversational conditions, and zoom/scribe_v1 jumps from #4 to #1 while the public leaderboard leader drops. Models tuned for clean read speech degrade disproportionately on actual conversations.
The practical consequence is that rank is no longer the deciding variable. License, language coverage, streaming support, and cost per audio-hour are. Cohere Transcribe has been downloaded over 620,000 times in the past month with runtime support across transformers, vLLM, mlx-audio for Apple Silicon, a Rust port, and a WebGPU build — but its model card warns there’s no automatic language detection, no timestamps, no diarization, and it’s eager to transcribe silence. IBM Granite Speech 4.1 offers fewer languages but adds bidirectional speech translation, keyword-list biasing for names and jargon, punctuation and truecasing. Qwen3-ASR-1.7B covers 52 languages and dialects, including 22 Chinese dialects, making it the obvious starting point for anything touching Mandarin.
Use the leaderboard to build a shortlist. Do not use it to pick a winner. The field has matured to the point where accuracy differences are smaller than the operational headaches of picking the wrong tool for your actual workload.
💡 Key Takeaways
- The Open ASR Leaderboard average is not a fixed quantity — models are scored on different subsets of test data, making direct numerical comparisons between published figures meaningless.
- MOSS-Transcribe-preview-2B was fine-tuned via reinforcement learning on the leaderboard's own training splits, meaning its score measures the benchmark rather than generalized transcription capability.
- When Appen's private conversational test sets are toggled on, zoom/scribe_v1 jumps from #4 to #1, proving models optimized for clean read speech degrade sharply on spontaneous audio.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.