AI Pulse by Inblix

Alibaba's Qwen TTS edges past rivals in quality, but crawls at 16 chars/second

The Decoder · Jul 21, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Alibaba's Qwen TTS edges past rivals in quality, but crawls at 16 chars/second

Alibaba just claimed the top spot on a major text-to-speech leaderboard with a model that sounds great — so long as you’re not in a hurry. Qwen-Audio-3.0-TTS-Plus hit an Elo score of 1,236 on Artificial Analysis’ Speech Arena, nudging past Simba 3.2’s 1,234. Google’s Gemini 3.1 Flash TTS and Cartesia’s Sonic 3.5 round out the chasing pack. The rankings make one thing clear: Alibaba can now build synthetic voices that people genuinely prefer in blind tests.

But the victory comes with a glaring asterisk. The Plus model generates speech at a glacial 16 characters per second. That’s barely half the speed of Simba 3.2 (30.2) and an order of magnitude slower than Sonic 3.5, which screams along at 120. For applications where latency matters — interactive voice agents, real-time translation, live narration — this is a non-starter. Alibaba does offer a Flash variant that responds in about 300 milliseconds, but that tradeoff between quality and speed is the whole ballgame right now.

What makes Qwen interesting isn’t just the Elo score. It supports 16 languages, including Tagalog, Malay, Thai, and Vietnamese — languages that most Western TTS providers treat as an afterthought. It also handles several Chinese dialects. Users can steer the emotional delivery with plain English instructions or drop in tags like “[angry]” or “[giggles]” for nonverbal texture. The model reportedly manages noisy reference audio better than its predecessors when cloning voices, which is the kind of practical improvement that matters more than leaderboard bragging rights.

At $27.60 per million characters through Alibaba Cloud Model Studio, it’s priced like a premium product with second-tier speed. The real question is whether Alibaba can close that throughput gap before the next generation of models from Google and Cartesia makes the quality lead irrelevant. Right now, you’re paying a premium for quality you have to wait for.

💡 Key Takeaways

  1. Qwen-Audio-3.0-TTS-Plus leads the Speech Arena leaderboard with an Elo of 1,236, but its 16 char/s speed makes it impractical for real-time applications where alternatives are 2-7x faster.
  2. The model supports 16 languages including Tagalog and Thai, plus several Chinese dialects — coverage that most Western TTS providers ignore entirely.
  3. Users can direct speaking style through natural language prompts or emotion tags like "[angry]" and "[giggles]", and the model handles noisy reference audio better than previous versions.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles