AWS released vLLM-Omni Deep Learning Container on SageMaker AI, allowing developers to generate images and videos from text prompts. It is the company's first major update to its SageMaker AI platform since the launch of bidirectional streaming for real-time speech in Part 1 of the series.

AWS reported the vLLM-Omni DLC packages tracked vLLM-Omni releases and adds routing middleware for SageMaker AI. That compares with previous versions of the DLC that focused solely on text generation without video capabilities.

vLLM-Omni is built on AWS's SageMaker AI platform and targets multi-modal applications that process or generate text, audio, images, and video through OpenAI-compatible APIs. Availability begins with the release of the vLLM-Omni DLC, initially for developers and researchers using AWS services.

"The vLLM-Omni DLC extends vLLM beyond text generation to models that process or generate text, audio, images, and video through OpenAI-compatible APIs," said the AWS blog post. The solution allows developers to deploy two endpoints from the same DLC image for image and video generation.

The announcement follows the release of Part 1 of the series, which covered real-time speech generation. AWS said the new feature expands the platform's capabilities to handle different models, payloads, and response patterns.

AWS did not say when the vLLM-Omni DLC will be available for other AWS regions, and the source raises the question of how the platform will scale to handle more complex video generation tasks. The sample includes a command-line workflow and a Streamlit application for developers.

Source: awsml