Alibaba's Qwen-Audio-3.0-TTS-Plus has claimed the top position on Artificial Analysis' Text-to-Speech Leaderboard for provider voices. The model achieved an Elo score of 1,236, surpassing Simba 3.2 with a score of 1,234. It is followed by Gemini 3.1 Flash TTS at 1,214 and Sonic 3.5 at 1,207. The model comes in two versions: Flash, designed for real-time interaction with about 300 milliseconds of latency, and Plus, focused on high-quality speech output. It supports 16 languages, including Tagalog, Malay, Thai, Vietnamese, and several Chinese dialects. Users can influence speaking style with natural language or add nonverbal cues using tags like '[angry]' or '[giggles]'.
Alibaba claims the model handles noisy or echo-heavy reference recordings more effectively than previous versions when cloning voices. However, speed remains a challenge, with the model processing 16 characters per second, significantly slower than Sonic 3.5's 120 and Simba 3.2's 30.2. The model is priced at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available for review.
Alibaba's Qwen-Audio-3.0-TTS-Plus was released as part of its ongoing efforts to enhance text-to-speech technology. The model's performance on the Artificial Analysis leaderboard highlights its competitive edge in the field.
Source: thedecoder