The model builds on Self-Flow, the company’s method for bringing multimodal generation and understanding into the same underlying system. Black Forest Labs increased the amount of training data and computing resources used for FLUX 3, training it on video, images and audio simultaneously.
The company sees FLUX 3 as a step toward what it calls real-world visual intelligence: systems that can interpret environments, anticipate changes and support actions in both digital and physical settings. Its initial applications are divided across video, image creation and robotics.
FLUX 3 Video can generate clips with native audio lasting as long as 20 seconds in one pass. It supports text-to-video creation, animation from an initial image, visual references, video transformation and the continuation of existing footage and sound.
Users can also define key moments for transitions, create multilingual dialogue and generate content in different aspect ratios and visual formats. Black Forest Labs says the model can move between styles including informal camera footage, animation and more cinematic output.
Additional capabilities include text rendering, animated graphics and the ability to connect individual clips into longer sequences. Visual references can be reused across scenes to help preserve characters and other recurring elements.
For its preliminary testing, Black Forest Labs generated 10-second videos at 720p with accompanying audio. The company cautioned that both the model and its evaluation system remain under development and that performance could change during early access.
In those evaluations, FLUX 3 was selected over Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%. It received preference rates of up to 69% against Grok Imagine Video, 60% against Kling v3 Pro, 59% against Happy Horse v1 and 57% against Happy Horse 1.1.
FLUX 3 was also chosen in 52% of comparisons against both Seedance 2.0 and Gemini Omni Flash. Black Forest Labs described all of the findings as early results rather than final benchmarks.
The company says the video model currently performs particularly well when producing facial expressions, matching sounds to physical events and handling multiple languages. It is also testing ways to combine generated clips into sequences lasting several minutes while maintaining visual continuity between scenes.
FLUX 3 Video is available through an early-access program.
The image component of FLUX 3 is designed to create and modify images across different styles, dimensions and resolutions. Black Forest Labs says testing conducted during training showed improvements over previous FLUX models, particularly when interpreting complicated instructions and generating written text.
The model can render text in several languages and produce a broad range of visual treatments. FLUX 3 Image is expected to enter early access in the coming weeks, with further development planned before its wider release.
Black Forest Labs is also extending the same architecture into action prediction. The company is pursuing one approach that builds action prediction directly into FLUX 3 and another that adapts the model’s video backbone into specialized systems using a smaller amount of task-specific training data.
Mimic Robotics was among the first partners to receive access. The companies developed FLUX-mimic, which combines the FLUX 3 foundation with Mimic’s work in robot learning, dexterous manipulation and production deployment.
That collaboration is intended to test whether a model trained to understand motion and physical change can also serve as a starting point for robotic actions. Black Forest Labs says the pretrained video system provides knowledge of dynamics that can be adapted to individual tasks.
The broader release will be organized into four product lines. FLUX 3 Video will provide video and audio creation and editing through APIs and private weight access. FLUX 3 Action and FLUX-mimic will make action-prediction capabilities available through selected research and commercial partners.
FLUX 3 Image will offer image generation and editing through APIs and private weights. FLUX 3 Dev is planned as an open-weight multimodal backbone covering image, video and audio creation as well as action prediction.
Black Forest Labs plans to introduce each capability after an early-access stage used for feedback, safety testing and rollout preparation. The company also intends to publish more technical information about the architecture.
The roadmap ultimately points toward a single model that combines perception, language and action prediction. Black Forest Labs says it is already developing the next generation while gradually making the current FLUX 3 capabilities available.
This analysis is based on reporting from Black Forest Labs.
Images courtesy of Black Forest Labs.
This article was generated with AI assistance and reviewed for accuracy and quality.