Black Forest Labs launches FLUX 3 unified multimodal model for image, video, audio, and robot action
Black Forest Labs released FLUX 3 on July 23, a unified multimodal model trained in a single architecture spanning image, video, audio, and action prediction. The model supports text-to-video and image-to-video generation, video-to-video transformation from reference clips, generative audio continuation, keyframe-to-video transitions, multilingual dialogue, and agentic chaining of clips into longer sequences. FLUX 3 Video is now available in early access, with a developer-weight version on the way. The unified architecture contrasts with prior modality-specific generators, consolidating image and video capabilities that previously required separate models.
The release includes FLUX-mimic, developed jointly with Mimic Robotics, which pairs FLUX 3's backbone with robot learning expertise to enable action prediction and control in physical systems. FLUX-mimic can be deployed on actual robots, demonstrating that FLUX 3 is learning a sufficient world model capable of driving dexterous manipulation and real-world factory tasks. The team also claims strong performance on visual styles (candid footage to animation), typography and design generation, and diverse aspect ratios extending beyond conventional cinematic formats.
FLUX 3's claim of unified training across modalities and direct robot deployment is notable because it demonstrates generative capability extending beyond traditional media to embodied control. The open developer build represents a competitive move against Sora, Grok Imagine, and Gemini Omni—the community is now verifying independent reproduction of frontier multimodal capabilities outside the largest labs. Architects evaluating video-plus-control stacks for robotics or agentic video pipelines should assess FLUX 3 alongside paid APIs.