On July 23, 2026, Black Forest Labs proudly announced the launch of its groundbreaking multimodal foundation model, FLUX 3. This innovative model adopts a unified architecture that facilitates the simultaneous learning of images, videos, and audio. It boasts the capability to generate synchronized audio-video content of up to 20 seconds in a single go. Additionally, FLUX 3 encompasses a range of functions, including text/image/video-to-video conversion, seamless keyframe transitions, and multilingual dialogue support. In a collaborative effort with Mimic Robotics, the model has been instrumental in developing the video motion model, Flux-mimic. Currently, FLUX 3 is undergoing rigorous testing at an Audi factory, marking a significant step forward in the realm of multimodal AI technology.
