Black Forest Labs has moved the FLUX name into robotics with FLUX 3 Action, a 7B world action model that jointly predicts future visual states and control actions. The model takes a camera frame and a text instruction, then produces the next two seconds of actions. The architecture is a diffusion transformer with video and action tokens in one sequence; the text instruction is encoded with a frozen Qwen3-VL-4B component.
The release includes checkpoints fine-tuned for the DROID dataset and the SO-101 robot arm, with both integrated into LeRobot. Black Forest Labs also publishes code and says the weights are released under the FLUX Kommunity License v1.0. In its announcement, the company reports a 42.92% success rate on the RoboLab-120 benchmark for the DROID-tuned model, above the comparison systems shown in its table. That result is useful context, but it remains a publisher-reported benchmark until reproduced under the same benchmark and harness conditions.
For Open Model Weights, the license is as important as the architecture. “Open weights” describes access to the model artifacts; it does not by itself answer what commercial uses, redistribution or derivative work are allowed. FLUX 3 Action should therefore enter the registry only after the exact repository artifacts and the dedicated license terms are checked field by field. The release also broadens what a model registry needs to represent: not every open-weight model produces text, and future records increasingly need precise task and modality descriptions.
What this changes for the evidence layer
High-priority new-model candidate. Because the release uses a dedicated license, commercial-use classification should be based on the exact license text rather than the phrase “open weights.”
Black Forest Labs / Hugging Face
This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.
Open source article ↗