Black Forest Labs opened early access to FLUX 3 on July 23, a foundation model trained at once on images, video, and audio rather than on each modality separately. The company that built its name on still-image generation is now pitching itself as something broader: a lab building a general model of how the physical world behaves. That repositioning, not any single benchmark number, is the real news in this launch.
The technical claim is that a single model learns more from three correlated signals than from one. Black Forest Labs argues that sound has to match a physical impact, motion has to obey mass, and a video’s next frame has to follow logically from the one before it. Training on all three at once, the company says, forces the model toward a shared representation of physical cause and effect rather than three separate, disconnected skills. The approach is built on Self-Flow, the company’s earlier method for aligning generation and understanding within one architecture, scaled up with more compute and data.
In practice, FLUX 3 generates images and matched video-audio clips of up to 20 seconds from a single prompt, and can extend a reference image or clip into a new scene while keeping a character consistent. Black Forest Labs also says the model can chain individual clips into longer, multi-shot sequences and render accurate text across multiple languages, both in still images and inside animated typography.
The company published preference rates against several rival video generators: FLUX 3 beat Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, with narrower margins against Kling v3 Pro, Seedance 2.0, and Gemini Omni Flash. Those figures come entirely from Black Forest Labs’s own early evaluations, run on the company’s own preferred clip length and resolution. No independent lab has reproduced them, and the announcement includes no outside benchmark results to check them against.
The more consequential part of the launch is the smallest: action prediction. Black Forest Labs says the same video backbone that generates a clip can be fine-tuned into a policy that predicts what a robot arm should do next, and it has already handed early access to mimic robotics, a partner building a combined video-action model called FLUX-mimic that the two say has been tested on production tasks at Audi. That is a real research partnership, not a shipped product. Black Forest Labs has not disclosed deployment numbers, task success rates, or a timeline for anything beyond the mimic collaboration, and the robotics claim is explicitly a bet on future capability rather than a description of what customers can buy today.
The framing matters because it puts Black Forest Labs in the same lane as Nvidia, whose Cosmos platform makes an almost identical wager: that a model trained to predict pixels and physics can be repurposed to predict robot actions. Nvidia backs that bet with chip revenue and a physical-AI ecosystem already selling into manufacturing. Black Forest Labs is starting from an image-generation business and a single robotics partner. The launch plan reflects that gap: video and image generation ship first through APIs and private weights, action prediction stays limited to selected partners, and an open-weight variant called FLUX 3 Dev is promised for later without a firm date.
Teams building on FLUX for content generation should treat the video and image releases as production-relevant once early access opens broadly, but should not budget for FLUX-driven robotics workflows until Black Forest Labs publishes task-level results beyond the Audi pilot.
Announced by Black Forest Labs on its company blog on July 23, 2026.