AI Pulse by Inblix

Flux 3 beats Luma 93% of the time in early tests, adds native audio

The Decoder · Jul 23, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Flux 3 beats Luma 93% of the time in early tests, adds native audio

Black Forest Labs just raised the stakes in AI video. Their new Flux 3 model doesn’t just generate silent clips — it creates videos up to 20 seconds long with native audio baked in. That’s a first for the German startup, and early (self-reported) benchmarks suggest it can hang with the best. In head-to-head tests using 10-second, 720p clips, human evaluators preferred Flux 3 over Luma Ray 3.2 a staggering 93% of the time. It also beat Runway Gen-4.5 in 77% of comparisons and Grok Imagine Video in 69%.

The margins get tighter against the top tier. Flux 3 only edged out Kling v3 Pro 60% of the time and essentially tied with Seedance 2.0 and Gemini Omni Flash at 52% each. Tying Seedance is noteworthy — that model has already found work in Hollywood. BFL calls these results preliminary and no independent testing exists yet, so don’t mistake a vendor’s own eval for gospel. But if those numbers hold, Flux 3 jumps into the front rank overnight.

The architecture is what makes this interesting. Flux 3 is a multimodal foundation model trained on images, video, and audio simultaneously. BFL’s argument is straightforward: no single data type captures reality completely. Images give you spatial structure, video shows change over time, and audio reveals the relationship between physical events and the sounds they make. Train on all three at once and each modality compensates for gaps in the others. The company calls their training method Self-Flow — a unified approach where one model learns to generate and understand content through a shared internal representation.

There’s a robotics angle too. BFL partnered with Mimic Robotics to build Flux-mimic, a video-action model already being tested on production tasks at Audi. It’s part of a broader push toward what BFL calls “real-world visual intelligence” — models that perceive, predict, and act. The company is rolling everything out in phases. Flux 3 Video is available now, Flux 3 Image hits early access within weeks, and action prediction stays partner-only for the moment. BFL also plans to release open-weight access under the name Flux 3 Dev, which should make the open-source crowd very happy.

💡 Key Takeaways

  1. Flux 3 generates video with native audio — a first for Black Forest Labs — making silent AI clips feel increasingly outdated.
  2. Self-reported benchmarks show Flux 3 decisively beating Luma and Runway, but it only ties with leaders like Seedance and Gemini Omni Flash.
  3. BFL is bundling video generation, image generation, and action prediction into one multimodal architecture trained on images, video, and audio together.
  4. Flux-mimic, a robotics spin-off of the same architecture, is already running production tests at Audi factories.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles