DeepMind released Gemini 3.8 text-to-speech models on September 23, 2026, saying it transforms voice generation from static presets into a dynamic creative studio. It is the company's first major text-to-speech update since the launch of Gemini 3.5 Live Translate.
DeepMind reported its models achieved a score of 71.4 on Hume AI’s Voice Design Benchmark and ranked #1 overall, with 60.8 in accent modeling. That compares with previous models that did not achieve such benchmark scores.
Gemini 3.8 Flash TTS is built on natural language prompting and targets high expressive audio creation for gaming, audiobooks, and interactive media. Availability begins today, initially for developers and enterprises through the Gemini API and Google AI Studio.
"These models enable creators, developers, and enterprises to create richer, more expressive audio experiences," said Alan Cowen, Director, Research Science. The models support a wide range of use cases, including long-form content and dual-speaker screenplay control.
The announcement follows the launch of Gemini 3.5 Live Translate and 3.5 Transcribe, expanding DeepMind’s audio capabilities. DeepMind did not say when full enterprise access will be available, and noted limitations in voice remixing features, which are coming soon.
Source: deepmind