Google Launches Gemini 3.5 Transcribe, Its Most Accurate Speech-to-Text Model Yet

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-08-28 | By: Carlos Montiel | Reading time: ~4 minutes

Google DeepMind launched Gemini 3.5 Transcribe on August 26, 2026, a speech-to-text model the company describes as its most accurate transcription to date, supporting more than 85 languages with a word error rate of just 2.6% in non-streaming mode.

What Sets This Model Apart

Unlike conventional speech recognition, Gemini 3.5 Transcribe doesn't just transcribe audio literally: it strips filler words ("um," "like"), cleans up mid-sentence self-corrections, and aims to capture the speaker's actual intent to deliver polished, formatted text, according to Google's official announcement. The company reports a word error rate of 4.0% in streaming mode and 2.6% in non-streaming mode, with a 70% improvement in time-to-final-transcript versus Chirp 3, its previous model.

Where It's Available

The model is starting to roll out in the Gemini app for macOS and Android's Rambler feature, and Google confirmed it will soon enable voice-to-text dictation in any Chrome web field. For developers, it's already available in public preview through Google AI Studio and Google Antigravity, enabling API integration into custom products.

Gemini 3.5 Transcribe — spec sheet Supported languages: 85+ Word error rate: 4.0% (streaming) / 2.6% (non-streaming) Speed improvement vs. Chirp 3: 70% faster to final transcript Availability: Gemini App (macOS), Rambler (Android), Chrome coming soon Developer access: Google AI Studio, Google Antigravity (public preview)
What this means for companies with voice support or sales: a transcription model under 3% error with native support for 85+ languages directly cuts down the human review work in call center transcription, sales meeting notes, and customer service quality analysis. For companies in Guatemala and Latin America running multilingual contact centers or serving markets across several countries in the region, integrating a model like this via API can turn hours of audio into structured, analyzable text nearly in real time, without relying on slower and pricier manual transcription services.
Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contact · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com