Black Forest Labs has released Flux 3, a multimodal foundation model that generates videos with native audio up to 20 seconds long. This marks the first time the company has integrated audio directly into video creation. The model supports text-to-video, image-to-video, and video-to-video generation, along with keyframe-based transitions and multilingual dialogue. According to the company, Flux 3 excels at capturing human facial expressions and aligning sounds with physical events. The release follows early testing where the model outperformed several competitors in video generation tasks. Source: thedecoder

In early evaluations using 10-second clips at 720p, Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrowed against stronger competitors, with Flux 3 being preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, and over Happy Horse 1.1 at 57 percent. BFL noted that the results are preliminary and no independent tests are available yet. The company also mentioned that matching leading systems like Seedance and Gemini Omni Flash would position Flux 3 among the top video models. Source: thedecoder

Flux 3 is based on Self-Flow, BFL's approach for teaching one model to generate and understand content simultaneously. A multimodal transformer uses dedicated components to convert images, video, and audio into a shared internal representation and then turn it back into outputs. The model's action component can be extended for new uses, providing a foundation for robotics applications. BFL claims this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and the model's understanding of the physical world. Source: thedecoder