Google has introduced Gemini 3.5 Transcribe, an AI model designed to enhance voice-to-text transcription by refining speech input and reducing errors. The model is already used in Gboard's 'Rambler' feature on Pixel 11 devices and is set to expand across the Google ecosystem. According to the company, Gemini 3.5 Transcribe is significantly faster and more accurate than its predecessor, Chirp 3, with a 70 percent improvement in transcription speed and a live-speech error rate of 5.5 percent. This marks a modest improvement over Chirp 3's 7.32 percent error rate. The model's enhancements include the ability to edit out filler words like 'ums' and 'uhs' and to refine text in real time based on user corrections and custom vocabulary. It supports 85 languages and can handle up to three speakers in pre-recorded audio. However, the AI's ability to rephrase speech may not always be appropriate, particularly in contexts requiring precise wording. Google plans to integrate Gemini 3.5 Transcribe into more devices and applications, including the Gemini app on macOS and the Chrome browser, with broader availability expected soon. Developers can already access the model through the Gemini API and AI Studio's build model. The feature is also available in Antigravity, where it provides full access to screen context and chat history with user permission. As the model becomes more widely available, it aims to streamline voice input for a variety of tasks, from writing emails to interacting with AI chatbots.

Source: arstechnica