FLUX 3 Video creates clips up to 20 seconds with native audio from text and images
Black Forest Labs
Black Forest Labs just launched FLUX 3 Video, making 20-second AI clips with built-in audio and dialogue. It claims to beat rivals like Seedance 2.0 in head-to-head tests, and an open-weight version is coming.
Based on reporting by Black Forest Labs — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Black Forest Labs, the outfit behind the FLUX image models, is now letting anyone with API access generate video. FLUX 3 Video went live today through the BFL API and a handful of partners, and the pitch is bigger than just another text-to-video tool: it wants to be a reality model, not a highlight-reel generator.
The specs are solid without being outrageous. Clips run up to 20 seconds, native resolution tops out at 720p HD with a Full HD upscale option, and — this is the part that separates it from a lot of the field — audio gets generated alongside the footage rather than bolted on afterward. Dialogue comes with lip-sync across more than a dozen languages, including English, Mandarin, Hindi, Punjabi and Turkish. Sound effects and ambient noise are baked into the same generation pass.
What's more interesting is the flexibility BFL is leaning into. You can feed it a start frame, an end frame, and let it fill the middle. You can hand it four seconds of an existing video and tell it to keep going, and it'll try to match camera motion and dialogue across the cut. It can also switch scenes and camera angles inside one single generation, which is the kind of multi-shot coherence that's tripped up earlier video models. There's a draft mode too, a cheap low-fidelity pass so you can lock in a concept before paying for the full-quality render — a genuinely practical touch for anyone burning credits on iteration.
On performance, BFL says internal human evaluations put FLUX 3 ahead of existing state-of-the-art models in text-to-video by a clear margin, and tied with Seedance 2.0 in image-to-video while beating everything else. Take that with the usual grain of salt reserved for a company grading its own homework, but the fact that they're naming Seedance 2.0 directly as the benchmark tells you where the real competition sits right now.
BFL also flagged that it ran the model through third-party safety testing with Cinder, specifically checking for non-consensual intimate imagery and CSAM risks before shipping. That's become close to standard practice for any lab releasing a video generator capable of realistic human likenesses, and for good reason given how quickly these tools get abused. Next up: reference-based generation using combinations of image, video and audio inputs, plus a FLUX 3 Image model and — notably — an open-weight FLUX 3 Dev variant still to come.
My take — AI-written commentary, not fact-checked reporting
An open-weight video model on the roadmap is the detail that actually matters here, not the 20-second clips or the lip-sync party trick. Closed APIs from well-funded labs are a dime a dozen this year; what's scarce is anyone shipping serious video generation weights the community can actually run, inspect, and fine-tune. If BFL follows through on FLUX 3 Dev the way it did with the original FLUX line, that's the story worth watching, not another leaderboard claim against Seedance.
Read more about this at: Black Forest Labs
Related stories
Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction
MarkTechPost · 1 month ago ·
15
[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model
Latent Space · 1 month ago ·
9