TLDRocket
Sign in

FLUX 3 Expands Black Forest Labs' Image Generation Capabilities

The Neuron Covered by 3 sources

Black Forest Labs just launched FLUX 3, a model that generates video, images, and audio together instead of separately. It's their bet that AI needs to learn sound, sight, and motion as one thing to actually understand the physical world.

Black Forest Labs, the German outfit behind the original FLUX image models, has opened early access to FLUX 3, and this one is a different animal from its predecessors. Rather than bolting audio onto a video model or treating images as their own island, FLUX 3 trains on video, images, and audio simultaneously inside a single architecture the company calls Self-Flow. The pitch is straightforward once you sit with it: a photo freezes a moment, a video adds motion, audio reveals impacts and causes vision can miss, and language ties all of it to intent. Train on one and you learn that slice. Train on all of them together and the modalities start correcting each other, so the sound has to match the collision and the motion has to respect mass.

The practical result is a model that can generate 20-second video clips with native audio baked in from the start, not layered on afterward. FLUX 3 handles text-to-video, image-to-video animation, video-to-video style transfer that keeps a character consistent across scenes, keyframe-controlled transitions, and even multilingual dialogue. Black Forest Labs says it's already stitching individual clips into multi-shot sequences lasting several minutes without characters drifting in appearance, which has been one of the more stubborn problems in generative video.

On the numbers, the company claims FLUX 3 Video beat Runway Gen-4.5 in 77% of head-to-head comparisons and Luma Ray 3.2 in a lopsided 93%. Against closer competitors like Kling v3 Pro, Seedance 2.0, and Google's Gemini Omni Flash, the margins shrink to somewhere between 52% and 60%. Those figures come from Black Forest Labs itself, using an evaluation harness the company admits is still under construction, so treat the specific percentages as a starting point rather than gospel. A separate image track, still in midtraining, reportedly handles complex prompts and multilingual text rendering noticeably better than earlier FLUX releases, with a public early access window coming in the following weeks.

The more interesting swing, though, is toward robotics. Black Forest Labs partnered with mimic robotics to build FLUX-mimic, which repurposes the video backbone as what they call a dynamics-aware foundation for training dexterous manipulation models with comparatively little task-specific data. Audi is apparently already testing it on production tasks. That's the real bet buried in this release: if a model genuinely learns how objects fall, sound, and move by watching enough video and audio together, that same representation should transfer to a robot arm figuring out how to grip something without crushing it. Content creation and physical AI, in this framing, aren't separate product lines. They're two applications of the same underlying world model.

Black Forest Labs is rolling this out in stages: video and audio APIs first, then image, then action prediction through select partners, and eventually an open-weight version called FLUX 3 Dev. The company frames all of it as one step toward unifying perception, action, and language prediction in a single model, which is a big claim to make off an early-access launch with self-reported benchmarks. Still, the architecture choice is the more defensible part of the story, independent of whose numbers you trust.

My take

I'll believe the robotics transfer story once someone outside Black Forest Labs and its handpicked partner mimic puts FLUX-mimic through independent testing, because self-reported win rates against your own harness are marketing, not evidence. That said, training one model across video, audio, and images instead of stapling separate systems together is the correct direction, and I'm glad they're promising an open-weight FLUX 3 Dev rather than locking the whole stack behind an API. Watch whether that open release actually ships on schedule — that's the real signal, not the 93% win rate against Luma.

Read more about this at: The Neuron

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.