Black Forest Labs Launches FLUX 3 Video AI Model With Native Audio and 20-Second HD Video Generation

Black Forest Labs Launches FLUX 3 Video AI Model With Native Audio and 20-Second HD Video Generation

Black Forest Labs has made an initial version of FLUX 3 Video generally available through the BFL API and select partners, bringing the company's multimodal video generation model to broader use for the first time. The release supports text-to-video and image-to-video generation, producing clips of up to 20 seconds in HD resolution with native audio, while Full HD output is available through upscaling.

The launch marks the first general availability of FLUX 3's video generation capabilities. Black Forest Labs describes FLUX 3 as its frontier multimodal model built to generate and predict video, audio, images and actions, with an architecture designed to model different forms of media together rather than as separate systems.

According to the company, FLUX 3 Video is designed to interpret both simple prompts and detailed instructions while maintaining coherent scene transitions, camera movements and audio. The model can generate multiple shots within a single clip, render typography as part of a scene, and produce spoken dialogue with natural accents and lip-syncing across multiple supported languages.

The initial release includes several generation modes. Users can create videos directly from text prompts, animate still images, define beginning and ending frames or multiple keyframes, and extend existing video clips by supplying up to four seconds of video and audio. The continuation feature carries movement, dialogue, camera behavior and sound across the transition between the original footage and newly generated content.

Black Forest Labs also introduced a Draft Mode aimed at speeding creative iteration. The feature generates lower-cost preview versions of a prompt before rendering the approved concept at full quality, while preserving the same subjects, composition and motion between the preview and final output.

The company said FLUX 3 Video supports HD (720p) generation and Full HD (1080p) output through upscaling. Native audio generation is included, allowing dialogue, ambient sound and sound effects to be created alongside video instead of being added separately.

Black Forest Labs said the model supports multiple languages, including English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi, with lip-syncing designed to match generated speech. The company also said FLUX 3 combines knowledge acquired during pretraining with real-time grounding to support applications such as documentaries and educational videos generated from short prompts.

In its evaluation of the model, Black Forest Labs said FLUX 3 Video delivers state-of-the-art performance for both text-to-video and image-to-video generation. The company reported that human evaluators preferred FLUX 3 over competing models in text-to-video generation, while its image-to-video performance tied Seedance 2.0 and exceeded other evaluated models.

The release also includes safeguards intended to reduce misuse. Black Forest Labs said it worked with third-party safety company Cinder to evaluate FLUX 3 Video before launch across risks including non-consensual intimate imagery and child sexual abuse material.

Looking ahead, the company said future updates will expand controllability and introduce additional multimodal capabilities, including video generation using combinations of image, video and audio references. Black Forest Labs also plans future releases of FLUX 3 Image for image generation and editing, along with FLUX 3 Dev as an open-weight variant.

This analysis is based on reporting from Black Forest Labs.

Image courtesy of Black Forest Labs.

This article was generated with AI assistance and reviewed for accuracy and quality.

Last updated: August 5, 2026

About this article: This article was generated with AI assistance and reviewed by our editorial team to ensure it follows our editorial standards for accuracy and independence. We maintain strict fact-checking protocols and cite all sources.

Word count: 539Reading time: 0 minutes
Browse All Articles
Share this article:
Next Article

AI News Daily

Breaking Intelligence • Since 2023

Join hundreds of thousands of AI professionals who start their day with our curated newsletter. Get breaking news, expert analysis, and exclusive insights.

Stay Ahead of AI

Get the latest AI breakthroughs, tools, and insights delivered to your inbox every week.

Free forever Unsubscribe anytime No spam guarantee

Go Premium

Unlock unlimited AI tools and an ad-free reading experience designed for AI professionals.

• Ad-free experience• Premium AI tools
Start Free Trial

14-day free trial • Cancel anytime
Plus $9/mo • Pro $90/yr (2 months free)

Follow Our Community

ChatAI

Breaking Intelligence

Your daily briefing on what matters in AI. Trusted by developers, researchers, executives, and AI enthusiasts worldwide.

© 2026 ChatAI. All rights reserved.