Google DeepMind launched Gemini 3.5 Transcribe on August 26, 2026, a speech-to-text model the company describes as its most accurate transcription to date, supporting more than 85 languages with a word error rate of just 2.6% in non-streaming mode.
Unlike conventional speech recognition, Gemini 3.5 Transcribe doesn't just transcribe audio literally: it strips filler words ("um," "like"), cleans up mid-sentence self-corrections, and aims to capture the speaker's actual intent to deliver polished, formatted text, according to Google's official announcement. The company reports a word error rate of 4.0% in streaming mode and 2.6% in non-streaming mode, with a 70% improvement in time-to-final-transcript versus Chirp 3, its previous model.
The model is starting to roll out in the Gemini app for macOS and Android's Rambler feature, and Google confirmed it will soon enable voice-to-text dictation in any Chrome web field. For developers, it's already available in public preview through Google AI Studio and Google Antigravity, enabling API integration into custom products.
Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.
Contact Carlos Montiel