Alibaba has released its latest video generation model, Wan3.0, which can produce AI videos up to 30 seconds long. The model accepts input in the form of text, PDFs, web pages, and PowerPoint files. This represents a doubling of the video length compared to its predecessor, Wan2.5. The model also suggests the optimal video length based on user prompts and includes a tool to extend existing videos. According to Alibaba, Wan3.0 processes text, images, video, and audio simultaneously. A single prompt can include up to ten images, five videos, and five audio clips. Web pages and documents like PDFs or PowerPoint presentations can also serve as input sources, transforming static data into dynamic video content. Source: thedecoder
Wan3.0 aims to address common issues in AI-generated videos, such as visual drift and distortion, particularly in faces and user interfaces. The model maintains consistency in details from reference materials, including characters, props, and spatial layouts. It is accessible through the wan.video website, Alibaba Cloud Model Studio, or via API on Qwen Cloud. Two pricing tiers are available: a Standard version currently at a 30 percent discount and a faster Prime version. Resolution options include 480p, 720p, and 1080p, with corresponding cost variations. Source: thedecoder
Alibaba is positioning Wan3.0 for a variety of applications, including film production, robotics training, and content creation. The model can help businesses generate marketing and training videos from text and images, while developers can use it to create realistic simulation footage for autonomous vehicles and robotics systems. The launch coincides with increased AI spending by Alibaba, which recently announced a major share sale to fund its AI initiatives and reported a 75 percent year-over-year drop in quarterly profit due to higher AI investments. Source: thedecoder