AWS released the Qwen3-TTS-12Hz-1.7B-Base model on Amazon SageMaker AI, allowing developers to generate speech in a target speaker’s voice without retraining a model. It is the company’s first text-to-speech model update since the introduction of the Qwen3-TTS family.

AWS reported the model supports 10 languages, including Chinese, English, Japanese, and Korean, with a base variant of 1.7 billion parameters. That compares with previous models that supported fewer languages and smaller parameter counts.

The Qwen3-TTS-12Hz-1.7B-Base is built on the Qwen3-TTS-Tokenizer-12Hz speech tokenizer and targets use cases such as content localization and customer engagement. Availability begins with the Amazon SageMaker JumpStart platform, initially for media teams, educators, and application developers.

"You can generate new speech in a target speaker’s voice from a short reference recording, without retraining a model," said the Qwen team at Alibaba Cloud. The model captures vocal characteristics such as timbre, pitch, and cadence to apply to new text.

The announcement follows the launch of the Qwen3-TTS-12Hz-1.7B-CustomVoice variant. AWS did not say how the model performs on cross-lingual cloning benchmarks, and the source raises the question of how well the model adapts to domain-specific speech patterns.

The model can be deployed on a fully managed real-time endpoint, with automatic scaling and health monitoring provided by Amazon SageMaker AI.

Source: awsml