According to the company, FLUX 3 Video is designed to interpret both simple prompts and detailed instructions while maintaining coherent scene transitions, camera movements and audio. The model can generate multiple shots within a single clip, render typography as part of a scene, and produce spoken dialogue with natural accents and lip-syncing across multiple supported languages.
The initial release includes several generation modes. Users can create videos directly from text prompts, animate still images, define beginning and ending frames or multiple keyframes, and extend existing video clips by supplying up to four seconds of video and audio. The continuation feature carries movement, dialogue, camera behavior and sound across the transition between the original footage and newly generated content.
Black Forest Labs also introduced a Draft Mode aimed at speeding creative iteration. The feature generates lower-cost preview versions of a prompt before rendering the approved concept at full quality, while preserving the same subjects, composition and motion between the preview and final output.
The company said FLUX 3 Video supports HD (720p) generation and Full HD (1080p) output through upscaling. Native audio generation is included, allowing dialogue, ambient sound and sound effects to be created alongside video instead of being added separately.
Black Forest Labs said the model supports multiple languages, including English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi, with lip-syncing designed to match generated speech. The company also said FLUX 3 combines knowledge acquired during pretraining with real-time grounding to support applications such as documentaries and educational videos generated from short prompts.
In its evaluation of the model, Black Forest Labs said FLUX 3 Video delivers state-of-the-art performance for both text-to-video and image-to-video generation. The company reported that human evaluators preferred FLUX 3 over competing models in text-to-video generation, while its image-to-video performance tied Seedance 2.0 and exceeded other evaluated models.
The release also includes safeguards intended to reduce misuse. Black Forest Labs said it worked with third-party safety company Cinder to evaluate FLUX 3 Video before launch across risks including non-consensual intimate imagery and child sexual abuse material.
Looking ahead, the company said future updates will expand controllability and introduce additional multimodal capabilities, including video generation using combinations of image, video and audio references. Black Forest Labs also plans future releases of FLUX 3 Image for image generation and editing, along with FLUX 3 Dev as an open-weight variant.
This analysis is based on reporting from Black Forest Labs.
Image courtesy of Black Forest Labs.
This article was generated with AI assistance and reviewed for accuracy and quality.