AWS released vLLM-Omni, a new Deep Learning Container, enabling real-time voice applications with bidirectional streaming. The update allows text-to-speech models to stream audio output as text is generated. It is the company's first major update to its SageMaker AI infrastructure since the release of vLLM in 2023.
AWS reported the vLLM-Omni DLC supports models that process and generate text, audio, images, and video, with a focus on streamed speech for real-time voice applications. The update compares with earlier vLLM versions, which were limited to text generation. The new DLC adds routing middleware for SageMaker AI and supports OpenAI-compatible APIs.
vLLM-Omni is built on AWS's SageMaker AI platform and targets real-time voice applications such as customer service assistants and accessibility tools. Availability begins with the release of the vLLM-Omni v1.5 DLC, initially for developers and enterprises using AWS services.
"You use SageMaker bidirectional streaming to send text and receive audio chunks over the persistent connection," said the AWS blog post. The post added that the vLLM-Omni DLC enables developers to deploy models that process and generate multiple modalities.
The announcement follows the release of the first part of a series on specialized DLCs for multimodal inference. AWS said the series pairs focused use cases with deployment examples and reproducible benchmarks where they add useful evidence.
AWS did not say how the vLLM-Omni DLC will perform with other models beyond TTS, and the open question or limitation the source raises is whether the DLC can scale to more complex multimodal applications. The source says the series will continue with Part 2 applying the vLLM-Omni DLC to image and video generation.
Source: awsml