TLDRocket
Sign in

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

MarkTechPost Asif Razzaq

MiniMax dropped H3, a video model that takes text, images, video and audio all at once and spits out 2K clips with real stereo sound. It's API-only for now, but it's already beating rivals on editing tasks while costing less per second.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

MiniMax has a habit of shipping fast and talking price, and H3 continues that pattern. Launched July 31, 2026, this isn't just another text-to-video generator with bolted-on features. MiniMax built it to treat text, images, video clips, and audio as one shared context, so you can say something like "use the camera move from Video 1, make the character from Image 2 sing, and sync the vocals to Audio 3" and get a coherent 2K clip back, four to fifteen seconds long, with native stereo audio baked in.

What makes this more than marketing is the plumbing underneath. MiniMax rebuilt its captioning system, called Contextual Omni Representation, so descriptions capture the relationship between source material and target output rather than just describing the target in isolation. That trims roughly 100,000 tokens of inference work down to about 4,000. Then there's H3-VAE, a new tokenizer that MiniMax says delivers a 4x jump in effective sequence length — which is the actual reason 2K output is affordable at all, not just possible. Pair that with a transformer architecture built specifically for H3, ditching the old Hailuo-02 base because multimodal inputs tripled the variance in sequence lengths the model had to handle, and MiniMax claims nearly 30% better training throughput as a result.

The detail I find most interesting is how H3 handles resolution. Instead of running a separate upscaler on low-res output — the standard approach, and one that tends to hallucinate fine detail — H3 regenerates its own output by re-reading the full multimodal context. That matters a lot if you're rendering a product label or a storefront sign and need the text to actually be legible rather than upscaled guesswork.

Right now you can only reach H3 through MiniMax's API or the Hailuo AI consumer app; open weights are promised "in the coming days" but haven't landed. The API itself is fairly constrained — up to nine reference images, three reference video clips capped at 15 seconds combined, three audio clips (which can't be submitted alone, only alongside video or images), and a 12-file total across all inputs. MiniMax says 2K generation costs under a third of what competitors charge per second, with 768p running under half the price of rival 720p offerings; third-party trackers cite $0.13 per second, or about $1.95 for a full 15-second 2K clip, though MiniMax's own pricing page hadn't caught up to list it at time of writing.

On actual quality, the picture is mixed. Artificial Analysis, per South China Morning Post's reporting, ranks H3 first in video editing but behind Google's Gemini Omni Flash in straight text-to-video, and behind both Gemini Omni Flash and Seedance 2.0 in image-to-video. So H3's real edge right now isn't raw generation quality — it's the unified, natural-language control over editing and reference tasks that used to require a pile of separate specialist models.

My take — AI-written commentary, not fact-checked reporting

The unified-context trick is the genuinely interesting part here, not the 2K resolution headline everyone's running with. Folding six separate video tasks into one natural-language interface is the kind of architectural move that actually changes workflows, while resolution bumps are just spec-sheet theater that every lab leapfrogs every few months. I'll believe the open-weights promise when I see a repo, not a press release — "coming days" from a Chinese lab racing to out-cheap Google and Seedance usually means "whenever it's competitively convenient."

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.