xAI has released its latest video model, Grok-Imagine-Video-1.5, which now includes support for text, image, and voice references. This update enables users to generate up to 1080p video without needing a starting image. The new features allow for text-to-video generation, where users can describe a shot and create video from that description. Additionally, the model supports native 1080p resolution for both text-to-video and image-to-video workflows, according to xAI. The update is now generally available on grok.com/imagine, iOS, and Android platforms.

Image and voice references are now available in the US for SuperGrok Heavy and SuperGrok Plus tiers, with broader rollout planned for other tiers in the coming days. Users can provide a character image and a voice reference to ensure consistency across scenes, maintaining the same face and voice throughout. The model allows up to seven references per generation, enabling users to keep a character and swap scenes, or keep a scene and swap characters. Voice reference support is also available through the API upon request.

The update includes improvements in motion, physics, and audio, as noted in xAI's announcement. The new version of the model is available through the xAI API, with voice reference support being an optional feature. Developers can access the model using the xAI SDK, as demonstrated in the provided code example. The company emphasized that these updates enhance the capabilities of its video generation tools, making them more versatile and user-friendly.

Source: xai