Black Forest Labs, the German AI company behind the FLUX image models, has opened FLUX 3 Video to developers with API access and a limited group of partner platforms, according to The Decoder. That is a narrower rollout than a full public launch: consumers cannot yet reach the model through a standalone app. The system produces clips in HD and Full HD resolution, running as long as 20 seconds, and generates matching audio in the same pass, covering spoken dialogue, incidental sound effects, and environmental noise, rather than dubbing a silent clip afterward.
BFL frames the release as a win over three named rivals. Its own testing places FLUX 3 first in Elo-style rankings for text-to-video, scoring 1,135, and for image-to-video, scoring 1,051, ahead of Seedance 2.0, Minimax H3, and Google’s Gemini Omni Flash. Those numbers come from evaluations that Black Forest Labs designed, ran, and picked the competitors for. A company topping a leaderboard it built itself is a marketing claim, not independent confirmation. AI Insiders reported an independently run video-model ranking days ago, a different kind of test than a company grading its own release.
The native audio is the more technically demanding spec here, ahead of resolution or the 20-second ceiling. Generating sound and picture in a single pass forces a model to lock lip movement, ambient cues, and dialogue timing to the video from the start. That is harder than composing a soundtrack over a finished clip afterward, which is how most video-generation tools still work. FLUX 3 adds a further constraint: it produces synchronized, lip-matched speech across upward of 14 languages, so mouth shapes and cadence must hold up even as phoneme timing shifts from one language to the next. Few released video models attempt synchronized dialogue in a single language, let alone more than a dozen.
BFL also lists text-to-video, image-to-video, keyframing, video continuation, and multi-camera scene generation within one clip, plus the ability to render on-screen typography and pull from broad world knowledge, useful for tasks like documentary-style footage.
Pricing scales with each second of output. Text-to-video and image-to-video generation cost $0.06 a second in draft-quality HD, climbing to $0.17 a second at full quality and $0.29 a second in Full HD. Video-to-video generation runs higher at every tier, from $0.12 a second in draft quality up to $0.53 a second in Full HD. Audio is bundled into every rate rather than billed separately.
Teams evaluating video-generation vendors should treat BFL’s Elo scores as a shortlist signal, not a verdict. Test the multilingual lip-sync claim against real footage before shifting production budget: a pipeline that only works cleanly in English is a different product than one that holds up across 14 languages.
This account is based on reporting by Matthias Bastian for The Decoder, published August 5, 2026.