AI Pulse by Inblix

Flux.1 beat Midjourney and DALL-E 3 in 2024 — here's the open-source shift it triggered

Hugging Face Blog · Jan 31, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Flux.1 beat Midjourney and DALL-E 3 in 2024 — here's the open-source shift it triggered

It’s almost jarring to remember that just a year ago, we were all laughing at AI’s inability to count fingers. The pace of change in creative AI has been brutal and brilliant. 2024 wasn’t just about better images; it was the year the architecture underneath everything got a gut renovation. The old workhorse, the U-Net, was largely shown the door in favor of the Diffusion Transformer (DiT) and flow matching objectives. Stability AI telegraphed the shift with Stable Diffusion 3, but it was HunyuanDiT that actually shipped the first open-source DiT model, cracking the door open for a flood of innovation.

The real bombshell, however, was Black Forest Labs’ Flux.1. The [dev] version didn’t just compete—it set a new state-of-the-art, leapfrogging proprietary titans like Midjourney v6.0 and DALL·E 3 (HD) on key benchmarks. That moment shattered a lingering assumption: that you needed a closed-source, API-gated model to get production-grade results. Suddenly, the highest-fidelity image generation was a thing you could run on your own hardware, and the community responded with the kind of frantic energy we haven’t seen since the original Stable Diffusion dropped over two years ago.

Yet here’s the weird wrinkle. While Flux.1’s raw image quality is undeniably superior, the fine-tuning ecosystem around it is still playing catch-up. The article points to a fascinating lag: the zero-shot personalization techniques that exploded in 2024—things like InstantID and IP-Adapter FaceID that let you clone a face from a single photo with no training—are still clinging to the older SDXL architecture. They work better there, despite SDXL’s lower base fidelity. Why? The community simply hasn’t cracked the semantic code of the DiT architecture yet. We don’t fully understand where identity, style, or composition live inside these new transformer-based models the way we painstakingly learned to manipulate the U-Net’s layers.

That knowledge gap is the bottleneck, and it’s what makes 2025 so unpredictable. The article rightly frames this as the next frontier: identifying the semantic roles of DiT components so that the brilliant personalization and control tools built on SDXL can be ported to Flux-level quality. The raw horsepower is already here. The manual just hasn’t been written yet. When it is—and I suspect we’ll see breakthroughs by summer—the open-source creative toolchain will enter a phase where the gap isn’t just closed, but inverted, leaving proprietary services scrambling to match the flexibility of a model you can truly dissect and rebuild.

💡 Key Takeaways

  1. Flux.1 [dev] surpassed closed-source leaders Midjourney v6.0 and DALL·E 3 on benchmarks, proving open-source image generation can claim the top spot in raw quality.
  2. The architectural shift from U-Net to Diffusion Transformer (DiT) is the biggest technical disruption since Stable Diffusion's debut, but the community still lacks the deep semantic understanding of DiT needed for fine-tuning.
  3. Best-in-class personalization tools like InstantID remain locked on the older SDXL architecture because researchers don't yet know how to manipulate identity and style within Flux's transformer-based design, making that knowledge gap the central challenge for 2025.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles