Google released Flash TTS and Flash-Lite TTS on September 23, 2026, saying they support voice creation from text descriptions and enable multilingual speech generation. It is the company's first major update to its text-to-speech models since the launch of Gemini 3.8.
Google reported the models lead in most categories of Hume AI's text-to-speech benchmark, measured on a standard set of evaluation criteria. That compares with earlier models, which did not provide such benchmark details.
Flash TTS is built on Gemini 3.8 and targets creative uses such as podcasts, audiobooks, and game characters. Availability begins through the Gemini API and Google AI Studio, initially for developers and content creators.
"A text prompt can define a voice's role, accent, and vocal traits across a wide range of languages and dialects," said Google. The models also offer a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English.
The announcement follows Google's expansion into AI voice generation with the release of Gemini 3.8. Google said the models are part of its broader strategy to enhance AI-driven content creation, adding no further judgment.
Google did not say how the models will handle multilingual voice cloning, and raised the open question of how well the voice cloning feature will work with non-native speakers. The models are rolling out through the Gemini API and Google AI Studio, with enterprise API access to follow.
Source: thedecoder