MiniMax's H3 is the first open-source model to top an AI video ranking
Curated by the Inblix editorial team
For the first time, an open-weight model sits at the top of a major AI video benchmark. China’s MiniMax released H3, a 33-billion-parameter multimodal beast that Artificial Analysis ranks number one for Video Editing, second for Text-to-Video, and third for Image-to-Video. It’s a genuine milestone in a field dominated by closed, API-gated systems like Runway and Pika.
The model is architecturally interesting because it doesn’t just stitch frames together. It processes text, images, video, and audio simultaneously, spitting out clips between four and 15 seconds with stereo sound. The model card specifies you can feed a single prompt up to nine reference images, three video clips, and three audio clips. That’s a lot of creative control handed to the user, assuming you know how to wield it.
But “open” comes with asterisks. MiniMax kept two critical components closed: the 2K resolution upscaling module and H3-Context-IR, a system that translates prompts and references into a structured intermediate format the model understands. Without it, running H3 locally in ComfyUI caps output at 768p, and you’re left manually prepping context using the company’s published guides. That’s a friction point most casual users won’t tolerate. The license adds another wrinkle—commercial use is only free for companies with under $20 million in revenue, a cap clearly aimed at startups while forcing bigger fish to negotiate.
The timing is tough. ByteDance dropped its closed-source Seedance 2.5 the same day, which generates 30-second clips with baked-in audio, double H3’s maximum length. H3’s release feels like a strategic half-step: weights are public, fine-tuning is possible on custom footage or characters, but the full pipeline remains gated. It’s a smart way to build a developer ecosystem without giving away the crown jewels, though I suspect the community will quickly reverse-engineer the missing context prep step. The real test is whether H3’s open fine-tuning capability lets it outpace Seedance’s longer, more polished output in real-world creative workflows.
💡 Key Takeaways
- Artificial Analysis now ranks an open-weight model—MiniMax H3—first for Video Editing, a first for the benchmark.
- H3 can ingest up to nine reference images, three video clips, and three audio clips in a single prompt to generate clips up to 15 seconds with stereo sound.
- Two components remain closed-source: the 2K resolution module and the prompt-translation system, capping local output at 768p and requiring manual context preparation.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.