China's MiniMax H3 has become the first open-source model to achieve the top position in an AI video ranking. According to Artificial Analysis, the 33-billion-parameter model leads in Video Editing, ranks second in Text-to-Video, and places third in Image-to-Video. The model processes text, images, video, and audio together, generating clips ranging from four to 15 seconds with stereo sound. A single prompt can include up to nine reference images, three video clips, and three audio clips, according to the model card.

MiniMax released the H3 video model weights, making it accessible to developers and researchers. However, two components remain closed: the 2K resolution module and the H3-Context-IR, which translates prompts and reference material into a structured intermediate format. Users can run H3 locally in ComfyUI, but the maximum resolution is limited to 768p. They must handle context preparation using MiniMax's published prompting guides. The open weights allow fine-tuning on custom footage, characters, or a specific visual style, though commercial use is restricted to companies with annual revenue under $20 million.

The release of MiniMax H3 coincided with the launch of ByteDance's closed Seedance 2.5, which generates 30-second clips with built-in audio. This development highlights the growing competition in the AI video generation space.

Source: thedecoder