Robotics Weekly

Black Forest Labs opens FLUX 3 Action, a 7B robot model

The image model lab says its 7B open weights model beats the previous best open model by 6.1 points while running up to 3.95x faster.

Black Forest Labs, the lab behind the FLUX image models, released FLUX 3 Action, an open weights 7B world action model (it predicts future video frames and robot actions at the same time). The company says it takes first place on the RoboLab benchmark, beating the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.

Black Forest Labs
@bfl_ai
X
It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.
Sep 23, 2026 · View on X

The usual complaint about world action models is latency. Predicting video is expensive, so most teams bolt a separate action head onto the video model to get speed back. Black Forest Labs says FLUX 3 Action does not need that trade. It still predicts video and actions together, plans more than twice as far ahead as the strongest open VLA (a model that turns camera images and text into robot actions), and runs faster per second of robot motion.

The numbers

Animesh Garg pulled out the head to head. FLUX 3 Action scores 42.9% on RoboLab-120 against 36.8% for Cosmos 3 Nano, which is 16B parameters, so less than half the size at a better score. Garg also flags the architecture detail, separate noise to signal transitions for video and for actions plus asymmetric classifier free guidance, which is how the model avoids splitting the action head off from video synthesis.

The part worth watching is what happens when you pair it with a planner. Garg says putting F3A under a high level reasoner cuts execution costs by 29% and runtimes by 40% versus pure reasoning, because the fast policy handles the motor primitives and the LLM only gets called to correct course. That is the same split several groups have been circling, a slow brain that thinks occasionally and a fast policy that moves.

What is actually out

Weights, code, the fine tuning recipe, benchmarks and reproducible examples. Black Forest Labs says teams can fine tune on their own demonstrations to get a policy for a specific robot and task, and that with NVIDIA it integrated the model natively into Hugging Face's LeRobot, with edge deployment on NVIDIA Jetson.

The lab also claims promising results beyond robotics, training task specific policies for simulated environments like gaming, vehicle control and computer use. No numbers were given for those.

Shown versus claimed. Every benchmark figure here comes from Black Forest Labs and from Garg reading its report, and no outside group has run the comparison yet. The difference from most model launches is that the weights are public, so anyone who cares can check.

Animesh Garg
@animesh_garg
X
pairing F3A with high-level reasoners cuts execution costs by 29% and runtimes by 40% over pure reasoning
Sep 23, 2026 · View on X

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.