MiniMax H3 folds 6 video models into one and ships 2K for $1.95 a clip
Curated by the Inblix editorial team
MiniMax just erased the boundary between video generation and video editing. Their new H3 model, live today, isn’t a text-to-video tool with extra features bolted on. It’s a single pretrained system that reads text, images, video, and audio as one unified context and outputs video with native stereo sound.
This is a genuine architectural shift. Previous stacks split the problem into separate expert models for text-to-video, image-to-video, subject reference, motion transfer, and editing. H3 collapses all of that into one paradigm where relationships are expressed in natural language. Their demo prompt tells the story: reference the camera movement from Video 1, have the character in Image 2 sing, and match the vocals to Audio 3. That’s not stitching outputs together—it’s one model reasoning across modalities simultaneously.
The enabler is a rebuilt VAE tokenizer with a compression ratio that delivers a claimed 4× gain in effective sequence length. That’s what makes native 2K output economically viable, cutting both training and inference costs. On top of that, MiniMax scrapped the old Hailuo-02 transformer architecture for a design that separates understanding and generation workloads, boosting end-to-end training throughput by nearly 30%. The most practical innovation might be in-context regeneration: instead of a super-resolution module that guesses at fine detail, the base model regenerates its own low-resolution output by re-reading the original multimodal context. For brand logos and product text, that’s the difference between legible and mush.
Pricing lands at a reported $0.13 per second of 2K video, or about $1.95 for a 15-second clip—though MiniMax’s own billing page hasn’t been updated yet. Third-party benchmarks from Artificial Analysis place H3 at the top for video editing tasks, while trailing Google’s Gemini Omni Flash in text-to-video and sitting behind both Seedance 2.0 and Gemini Omni Flash in image-to-video. Open weights are promised “in the coming days,” but for now, the API is the only route in. For advertising variant generation, product videos, and game cinematics, the pitch is straightforward: one API call replaces a pipeline of half a dozen specialized tools.
💡 Key Takeaways
- H3 unifies six previously separate expert models into a single pretrained system that reads text, images, video, and audio as one context.
- A redesigned VAE tokenizer with 4× effective sequence-length gain is what makes native 2K generation economically viable.
- In-context regeneration replaces conventional super-resolution, preserving small text and brand details that upscalers typically hallucinate.
- Artificial Analysis ranks H3 first in video editing but behind Google and ByteDance models in text-to-video and image-to-video tasks.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.