FLUX 3 beats Runway 77% of the time, ships video and audio from one model
Curated by the Inblix editorial team
Black Forest Labs just dropped FLUX 3, and for the first time, a single set of weights handles images, video, and audio. No more stitching together separate models for sound and motion. The result is a model that generates 20-second clips with native, synced audio in one shot. It’s not just a video generator with sound tacked on. The research team’s core bet is that training on all modalities at once forces them to constrain each other — the sound has to match the impact, the motion has to obey the physics.
The numbers from BFL’s human preference tests are eye-opening. When people compared 10-second, 720p text-to-video clips with audio, they picked FLUX 3 over Luma Ray 3.2 an astonishing 93% of the time. Against Runway Gen-4.5, it won 77% of comparisons. The gap narrowed against heavy hitters like Grok Imagine Video (69%) and Kling v3 Pro (60%), and it’s basically a coin flip versus Seedance 2.0 and Gemini Omni Flash at 52%. Those are serious benchmarks.
Under the hood, FLUX 3 runs on Self-Flow, BFL’s method for aligning multimodal generation with a flow matching objective. The twist is the training budget: video prediction consumed over 95% of the compute, while audio represents less than 0.5% of the tokens. That imbalance tells you where the heavy lifting is. The same backbone also powers FLUX-mimic, a separate robot policy that runs at under 80 milliseconds on a single RTX 5090.
Access is tightly controlled. Video and Action modes are in early access now. Image generation follows later, and open weights come last. If you’re hoping to download this and run it locally tomorrow, you’re out of luck. But the trajectory is clear — BFL is building a unified model that sees, hears, and eventually acts, all from one architecture.
💡 Key Takeaways
- FLUX 3 generates video and native audio from one architecture, not separate models stitched together, and wins 77% of human preference tests against Runway Gen-4.5.
- Video training consumed over 95% of compute while audio used less than 0.5% of tokens, revealing where the real processing demands are in multimodal generation.
- The same Self-Flow backbone runs a robot policy at under 80ms on a single RTX 5090, hinting at BFL's ambition to unify perception, generation, and action.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.