TLDRocket
Sign in

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

MarkTechPost Asif Razzaq

Black Forest Labs just put out FLUX 3 Action, a 7B robot policy that tops RoboLab-120. It also comes with license limits, so this isn’t an open door for everyone.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Black Forest Labs, best known for the FLUX image models, has moved into robot control with FLUX 3 Action. The new model is a 7B open-weights World Action Model, and on RoboLab-120 it takes the top spot with a 42.92% task success rate.

The appeal is obvious. Robot policies usually make you choose between richer world modeling and speed. FLUX 3 Action tries to keep both: it reads camera frames, robot state, and a text instruction, then predicts future video frames and the next chunk of actions together. That puts it in the same broad family as NVIDIA’s Cosmos 3 Nano, but BFL says it closes the speed gap with a smaller backbone and distillation.

The training mix is doing a lot of the work here. BFL says pretraining used image, video, and audio data, with video making up more than 95% of the tokens. Midtraining then blended pretraining data with action-aligned video, and that action data included game recordings, egocentric human hand video, handheld grippers, and teleoperation across 14 embodiments. Most robot data was mapped into a shared 50-dimension end-effector action space called EE50.

The numbers suggest pretraining mattered a lot. Without it, DROID-only training stayed below 1% on RoboLab. With pretraining, the same setup reached 11.6%. On the benchmark itself, FLUX 3 Action’s 42.92% beats Cosmos 3 Nano’s 36.8% by 6.1 points, while using 56% fewer parameters. BFL’s own multi-seed mean for the guidance-distilled FP8 checkpoint was 42.24% ± 0.36.

Hardware results point in the same direction. On a blind evaluation run by Positronic Robotics with a Franka arm, FLUX 3 Action completed 28 of 30 attempts, versus 27 for Cosmos 3 Nano, 20 for DreamZero, and 13 for π0.5. BFL also ships three checkpoints for the DROID policy, with the guidance-distilled version running faster than the base recipe and the step-distilled one going faster still, though with lower success. The catch is practical, not theoretical: the model can run on 24 GB cards with FP8 and text encoder offload, but the FLUX Kommunity License keeps it in non-commercial territory.

There’s also a hybrid setup with GPT 6 Astra, where the reasoner can execute, edit, or replace the policy’s predicted actions. BFL says that combination solved 90% of episodes at lower cost and lower time per success than pure Astra at maximum effort. So yes, this is another reminder that the robot stack is becoming a game of clever tradeoffs, not just bigger models.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of AI story: a model that actually has to survive contact with hardware, not just a leaderboard. But the non-commercial license means the industry gets another shiny robot brain to admire from behind the rope line. Open-weights, closed door — the modern classic.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.