Black Forest Labs FLUX 3: Multimodal Video + Audio + Action in One Model; Beats Seedance, Runway
Black Forest Labs unveiled FLUX 3 on July 23, 2026, a multimodal frontier model jointly trained across images, video, audio, and action prediction within a single architecture. FLUX 3 Video generates 20-second clips with native, in-sync audio; FLUX 3 Action predicts robotic movements; FLUX 3 Image improves still-image generation. The model builds on Black Forest Labs' Self-Flow approach for aligning multimodal understanding and generation, scaled to billions of parameters across multiple modalities simultaneously.
In early human preference evaluations, FLUX 3 Video beats Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, while performing at parity with Seedance 2.0 and Gemini Omni (52% wins). FLUX 3 is particularly strong on human facial expressions, multilingual text rendering, and associating sounds with physical events. FLUX-mimic, a video-action derivative, is already undergoing testing and deployment at Audi for general-purpose robotic manipulation, learning new tasks from as little as 30 minutes of robot data.
The launch is phased: FLUX 3 Video and Action are in early access now; FLUX 3 Image rolls out in coming weeks; open-weight FLUX 3 Dev is promised for later in 2026. No pricing has been disclosed. For builders of creative systems, robotics, and simulation stacks, the architectural unification—one model for video, image, audio, and action—signals a shift away from assembling modality-specific tools toward unified world models. FLUX.1 achieved open adoption in diffusion; FLUX 3's open-weight promise could accelerate custom robotics and agentic systems adoption.
Sources
- Primary source
- globenewswire.com
“FLUX 3 jointly learns from images, video, and audio within a unified architecture, and can also be extended to predict actions”
- venturebeat.com
“FLUX 3 Video already leads in early evaluations against frontier video models, and is particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities”
- tech.yahoo.com
“human reviewers preferred FLUX 3's output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%”