The Freiburg-based company, which most people know as an image generator developer, has announced the FLUX 3 model. It can create videos up to twenty seconds long with audio generated in the same pass as the visuals, can create and edit images, and the same foundation powers software that is currently controlling robots on Audi's production line.
The model understands physics
FLUX 3's strongest capability is video. The model creates clips up to twenty seconds long and generates audio for them at the same time. The model offers a wide range of modes: text-to-video, image-to-video (either using the image as the first frame of the animation or merely as a visual reference), video-to-video, where the model can carry over the same character from a reference clip into a different scene, extrapolation of continuing visuals and audio from an input recording, and transitions between predefined frames.
It can also handle dialogue in various languages, render text and animated typography, and accommodate unconventional aspect ratios. Individual clips can be linked together agentically to create multi-shot sequences lasting several minutes, while visual references help maintain the same character across scenes. According to the developers, the model performs best at rendering human facial expressions and matching audio to what is physically happening in the shot.
More interesting than the list of features is how the model was created. It was trained on images, video, and audio simultaneously within a single architecture, with video prediction consuming more than 95 percent of all computing power. To generate video that does not look strange, the model must learn about contact between bodies, movement, mass, and cause and effect. When it gets these wrong, people notice immediately. Audio is an easy discipline by comparison, accounting for less than half a percent of all tokens. The foundation is the Self-Flow method, which the company developed earlier and which connects content generation with an understanding of what is depicted. FLUX 3 is a significantly scaled-up version of it, trained on tens of millions of hours of general video and hundreds of thousands of hours of footage focused on how humans and robots manipulate objects.
How it compares with the competition
Black Forest Labs has released the results of preference testing in which people compared ten-second, 720p clips with audio generated from text. According to the company, FLUX 3 beat Luma Ray 3.2 in 93 percent of comparisons and Runway Gen-4.5 in 77 percent. It scored as high as 69 percent against Grok Imagine Video, 60 percent against Kling v3 Pro, 59 percent against Happy Horse v1, and 57 percent against version 1.1. The closest results were against Seedance 2.0 and Gemini Omni Flash, where it finished at 52 percent.
The company itself says these are preliminary figures. However, VentureBeat points out several things entirely missing from the announcement: the evaluation methodology, sample size, number of evaluators, pricing, service guarantees, and any benchmark for the image model. Without these, the company's in-house figures are difficult to verify. Competition in video generation also revolves around image quality, prompt adherence, character consistency, and audio, with new versions released within weeks, so rankings from one week may no longer apply the next.
The image component of the model will not be opened to the public for another few weeks. Based on the progress made during training, the company reports a significant improvement over older FLUX models in two areas in particular: processing complex instructions and rendering text in images, including in languages other than English.
FLUX-mimic, a robotics offshoot that extracts movements from the model
Black Forest Labs believes that a model capable of convincingly predicting video must have developed an internal representation of how the world works. And if that representation exists, it should also be possible to extract information from it about how to act in the world.
That is the purpose of the FLUX-mimic model, developed together with Mimic robotics. A lightweight action decoder was added to the FLUX 3 model's foundation, taking intermediate layers from the video prediction component and translating them into robot movements. The company also tested a second approach: integrating action prediction directly into FLUX 3. Human ratings of the generated clips initially fell by as much as ten percent, but after 3,500 steps the model had returned to its original level while also predicting actions. In benchmarks, the decoder outperforms existing vision-language-action models even when FLUX 3 itself is left entirely untouched and only the small decoder added on top is trained. Older approaches fail when similarly stripped down because they do not contain an understanding of real-world physics. When both the backbone and decoder are fine-tuned together, FLUX-mimic reports state-of-the-art success rates.
The goal is general-purpose object manipulation. A robot should understand what it sees, estimate what will happen if it pushes an object, and learn a new task from a minimal number of demonstrations. The figure attracting attention is this: for some tasks, less than half an hour of robot data is sufficient, whereas previous methods required thirty hours or more. “The hardest part of robotics is data,” says Elvis Nava, CTO of Mimic robotics. According to him, each new task normally requires hours of the robot repeating the same thing, while FLUX-mimic masters it within minutes. A side effect is also worth mentioning: if the robot fails to grasp an object, it tries again and completes the task without anyone having demonstrated that behavior.
The model is running on Audi's assembly floor
Audi is among the most highly automated manufacturers in the automotive industry, giving it a precise understanding of where conventional automation still falls short. Tasks involving flexible parts and delicate manipulation remain manual mainly for economic reasons: in premium manufacturing with many variants, reconfiguring a robotic cell for every case is too expensive.
Mimic deployed FLUX-mimic to arrange parts in divided pallets, insert control units into tight fixtures, assemble components, and handle soft materials such as seals and cables—tasks beyond the reach of conventional automation. “We have seen these robots solve complex manipulation tasks involving soft, flexible parts that would simply be impossible with conventional robotics,” says Christoph Schneider of the Audi Production Lab.
On a production line, however, it is not enough to respond correctly; the response must also come in time.The path from sensor input to an internal representation of the world takes less than 80 milliseconds on a single NVIDIA RTX 5090 graphics card. Mimic also optimized the rest of the pipeline, so the entire standalone robotic system responds in about 101 milliseconds, comparable to the response time of human vision.
Availability
FLUX 3 is divided into four product lines. FLUX 3 Video with audio and FLUX 3 Action have been in closed preview since last week. Anyone can apply, but access must be approved by the company. FLUX 3 Image is expected to launch within a few weeks, followed by general availability. The final component is FLUX 3 Dev, offering open weights for the multimodal backbone. For now, action prediction will remain limited to selected research and commercial partners, the first of which is Mimic robotics. There is currently no public API.



