Black Forest Labs has launched FLUX 3 Video, a model that generates HD and Full HD clips up to 20 seconds long. The model produces videos with native audio, including dialogue, sound effects, and ambient noise. It supports various video generation modes, such as text-to-video, image-to-video, and video continuation. FLUX 3 can also render typography directly in scenes and follow complex prompts using world knowledge for applications like documentaries. It generates lip-synced dialogue in more than 14 languages. Pricing is based on the length of video output, with different rates for draft mode and full quality. Draft mode for HD costs $0.06 per second for text-to-video or image-to-video, and $0.12 for video-to-video. At full quality, HD costs $0.17 per second for text-to-video or image-to-video, and $0.41 for video-to-video. Full HD costs $0.29 and $0.53 per second, respectively. Audio is included in all outputs. More video examples are available on the BFL blog.
BFL's internal tests ranked FLUX 3 at the top for text-to-video with an Elo score of 1,135 and image-to-video with 1,051. It outperformed Gemini Omni Flash, Minimax H3, and Seedance 2.0. The company claims FLUX 3 can render typography directly in scenes, follow complex prompts, and draw on world knowledge for uses such as documentaries. It also generates lip-synced dialogue in more than 14 languages. The model supports text-to-video, image-to-video, keyframes, video continuation, and multiple scenes and camera angles within a single clip.
According to the source, Black Forest Labs made FLUX 3 Video generally available through the BFL API and select partners. The model generates HD and Full HD clips up to 20 seconds long, with native audio that includes dialogue, sound effects, and ambient noise. It supports text-to-video, image-to-video, keyframes, video continuation, and multiple scenes and camera angles within a single clip. BFL says FLUX 3 can render typography directly in scenes, follow complex prompts, and draw on world knowledge for uses such as documentaries. It also generates lip-synced dialogue in more than 14 languages.
Source: thedecoder