TLDRocket
Sign in

Multimodal Models

51 summarised stories about Multimodal Models, each linking back to the original source. Browse all topics →

+ Follow this topic

Saturday, 1 August 2026

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

MarkTechPost 4 weeks ago 24

MiniMax released MiniMax H3, a multimodal video generation model that unifies text, image, video, and audio inputs into a single system rather than separate specialist models. The model generates 2K video clips lasting 4-15 seconds with native stereo audio at an estimated cost of $0.13 per second ($1.95 for a 15-second clip). The unification allows users to specify complex creative operations like referencing camera movement from one video and character actions from another using natural language, replacing traditional split pipelines across advertising, e-commerce, and film production workflows.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.