FLUX 3: from video generation to robot control
The same multimodal backbone that generates 20-second video with audio is being tested on industrial manipulation at Audi Production Lab.
In Audi’s Production Lab, a robot arm seats an electronic control unit into a tight-fitting fixture. When it misses a grasp, it corrects itself, grasps again, and completes the task. The arm and its control stack come from mimic, a Zurich-based robotics company building video-action models for industrial manipulation. At the center of the system is the backbone of a video generator: FLUX 3, the new multimodal foundation model from Black Forest Labs.
FLUX 3 was trained to predict how scenes evolve. Doing that well requires it to learn and represent how objects persist, how materials deform, and what follows when one object makes contact with another. FLUX-mimic tests whether those internal representations can also support robot control in the real world.
Black Forest Labs, an Air Street portfolio company, released FLUX 3 last week. The model learns jointly from images, video, and audio within a single architecture. Each modality supplies information the others lack: images capture spatial structure, video adds time and motion, and audio helps locate events such as speech and impact. BFL is extending the same underlying model to action prediction, offered through selected research and commercial partners.
mimic released FLUX-mimic alongside it. The system combines FLUX 3’s video backbone with mimic’s robot manipulation data, action decoder, hardware, and deployment stack. FLUX 3 Video is in early access now, while image generation and editing will follow soon, with action prediction available through selected partners and an open-weight multimodal backbone also planned. This model family will add more depth to BFL’s open weight track record of more than half a billion downloads to date.
One model for sight, sound, and action
New video models are commonly evaluated through pairwise human preference tests: raters watch two clips generated from the same prompt and choose the one they prefer. In BFL’s preliminary evaluations of 10-second, 720p text-to-video clips with audio, FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and Grok Imagine Video in up to 69%, with narrower margins against the strongest systems: 60% over Kling v3 Pro and 52% over Seedance 2.0 and Gemini Omni Flash.
The model can generate up to 20 seconds of video with audio in a single pass. It works from text, images, or video; carries characters and other visual references across multi-shot sequences; renders typography; and produces synchronized dialogue in several languages.
Producing a convincing video requires an internal account of the scene. Objects must persist when they leave the frame. Cables must bend where hands grip them. A part clicking into a housing should sound different from one dropped on the floor, and the sound must coincide with the visible impact. Joint training lets each modality constrain the others:
BFL trained FLUX 3 on tens of millions of hours of general video, including hundreds of thousands of hours focused on human and robot manipulation. Generation quality offers one view of what the model has learned. The robotics experiments test whether that knowledge is accessible enough to guide action.
From prediction to action
FLUX-mimic places a lightweight action decoder on intermediate features extracted from FLUX 3’s video prediction path. The backbone forms an expectation of how the scene will evolve, while the decoder translates that representation into chunks of robot actions. It never generates video at inference time, so a chunk of actions costs a single forward pass through the backbone rather than a full video rollout.
mimic benchmarked a preview version on a soft-body kitting task using a real robot. The company reports a 95% success rate without single-task fine-tuning or post-training, compared with 55% for an adapted π0.5 model trained on the same data mix and 70% for a flow-matching policy heavily post-trained on that task.
A separate ablation is more revealing about the representation itself. With the FLUX backbone frozen, FLUX-mimic remained competitive, while the π0.5 action decoder recorded no successful task completions. The result suggests that useful information about manipulation was already present in the pretrained video backbone, before it was adapted to the specific task.
The system also has to respond quickly enough to control a physical machine. BFL says the FLUX-mimic backbone can be optimized to run from input to internal representation in less than 80 milliseconds on a single NVIDIA RTX 5090. Further optimization across the action decoder, the inter-process path between sensors, model, and actuators, and real-time chunking that overlaps prediction with execution brings the full system’s reaction time to 101 milliseconds.
Audi Production Lab has been testing and deploying the system on real production use cases. These include kitting parts into structured trays, inserting electronic components into tight-fitting fixtures, assembling components, and manipulating flexible materials such as seals and cables, which mimic describes as materials conventional automation has never been able to handle.
“We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics,” said Christoph Schneider of Audi Production Lab.
The road to physical intelligence
Regular readers will recognize the thesis. When Odyssey raised its $310M Series B last month, we wrote that the next advance in machine intelligence may come from systems that build worlds, act inside them, and learn how reality behaves. FLUX-mimic brings that argument onto a factory floor: a representation learned through video generation is being used to control a robot performing industrial manipulation.
We see the other side of this via the flow of new investment opportunities at Air Street. Of the two dozen robotics companies we have seen this year alone, half are building manipulation policies, and they have largely converged on the same VLA-style architectures. But, the differences come down to data. Teleoperated demonstrations get collected task by task, they are slow and expensive to produce, and they transfer poorly when the task changes. A backbone that already encodes how objects behave changes what a small team has to pay for: the world knowledge comes from video that already exists, and the demonstration budget goes to the last mile. That’s where the frozen-backbone result shines, and being open-weights means that it can empower an entire ecosystem of robotics builders.
FLUX 3 Video is available in early access now. Image generation and editing will follow, alongside partner access for action prediction and an open-weight version of the multimodal backbone.









