Google has launched Gemini 3.5 Transcribe, a speech-to-text model designed for real-time transcription. The model recognizes over 85 languages and automatically corrects verbal stumbles, such as filler words like 'um.' It also formats the output text without requiring additional input. According to Google, the model achieves a word error rate of 4.0 percent for streaming and 2.6 percent for recorded audio.
It also boasts 70 percent lower latency compared to its predecessor, Chirp 3. The model uses 'function calling' to delegate tasks like image generation or web searches to other Gemini models. It comes with two interfaces: the Live API for real-time streaming with low latency and the Interactions API for processed audio with speaker attribution and timestamps.
The model is now available in Google AI Studio and on the Gemini Enterprise Agent Platform. It is already integrated into Gboard for Android and the Gemini app on macOS, with Chrome support expected soon. Source: thedecoder