FLUX 3 Video launches with lip-synced dialogue in 14 languages, tops Elo charts
Curated by the Inblix editorial team
Black Forest Labs just dropped FLUX 3 Video into general availability, and the early numbers are loud. The model sits at the top of BFL’s internal Elo rankings for text-to-video with a score of 1,135 — ahead of Seedance 2.0, Minimax H3, and Google’s Gemini Omni Flash. For image-to-video, it scores 1,051. Those aren’t marginal leads. They suggest a genuine shift in what’s possible when you bake native audio directly into the generation pipeline rather than bolting it on afterward.
What makes FLUX 3 different is the audio. We’re talking lip-synced dialogue in more than 14 languages, plus sound effects and ambient noise — all generated alongside the visuals. Most video models still treat sound as an afterthought, something you add in post. BFL is making the case that coherent audiovisual output from a single prompt is the new baseline. The model handles HD and Full HD clips up to 20 seconds, supports text-to-video, image-to-video, keyframes, video continuation, and can juggle multiple scenes and camera angles within one clip. It also renders typography directly in-scene, which is a niche but genuinely hard problem that trips up most diffusion-based systems.
Pricing is per-second and tiered by quality. Draft mode in HD runs $0.06 per second for text-to-video or image-to-video, doubling to $0.12 for video-to-video. Full-quality HD goes for $0.17 and $0.41 respectively, while Full HD jumps to $0.29 and $0.53. Audio is included at every tier, which takes the sting out of the cost if you were planning to license sound effects separately. For a 20-second Full HD clip with audio, you’re looking at roughly $5.80 — not cheap, but competitive with what Runway and Pika charge for shorter, silent outputs.
Here’s where I’m cautious: BFL’s Elo scores come from their own testing, and every lab’s internal benchmarks make their model look like a breakthrough. The real test is whether FLUX 3 holds up when users start stress-testing it on prompts the training data never anticipated — the kind of weird, specific, multi-character scenes where most video models fall apart. Still, the native audio integration is a structural advantage. If BFL trained the visual and audio components jointly, they’ve solved a synchronization problem that’s plagued the space since the first text-to-video demos. That alone makes this worth a close look.
💡 Key Takeaways
- FLUX 3 Video generates lip-synced dialogue in 14+ languages with native sound effects — audio isn't an add-on, it's part of the core generation.
- BFL's internal testing places FLUX 3 at the top of Elo rankings for both text-to-video (1,135) and image-to-video (1,051), surpassing Seedance 2.0 and Gemini Omni Flash.
- Full-quality HD pricing starts at $0.17 per second with audio included, positioning FLUX 3 competitively against silent outputs from Runway and Pika.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.