Google has announced 'Gemini 3.5 Transcribe,' a speech recognition model that can transcribe speech in real time while automatically removing filler words like 'uh' and 'um.'

Google has announced Gemini 3.5 Transcribe , a speech recognition model that can generate accurate, sophisticated, and formatted text from raw audio data without being hampered by noise, complex jargon, or spoken language.
Introducing Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is an AI transcription tool designed to seamlessly integrate into any development workflow, including voice agents, real-time captioning tools, and call content analysis pipelines. It is available via two APIs: the Live API and the Interactions API.
The Live API provides continuous two-way streaming with sub-second latency for interactive voice applications. The Interactions API identifies each speaker in recorded audio and transcribes it with word-by-word timestamps.
While users are already benefiting from Gemini 3.5 Transcribe through Android's voice input feature ' Rambler ' and the voice functionality of the macOS version of Gemini, developers will now be able to leverage Gemini 3.5 Transcribe to build similar functionalities using Google AI Studio's Gemini API and the Gemini Enterprise Agent Platform.
The transcription automatically removes meaningless interjections like 'uh' and 'um,' which are simply responses or pauses made while thinking about the next word. It achieves an average Word Error Rate (WER) of 4.0% during streaming and 2.6% when not streaming. It supports over 85 languages and also allows for the addition of custom vocabulary to accommodate specialized terminology and unique spellings.
Compared to the preceding transcription model 'Chirp 3,' it represents a significant improvement, and measurements using Artificial Analysis show that the time required to complete the final transcription has been reduced by 70%.
The following are streaming benchmark scores for FLEURS, a multilingual speech dataset for speech recognition and language evaluation models in 102 languages. Lower values indicate better performance, and Gemini 3.5 Transcribe is rated as superior to other models, including Chirp 3.

The same applies when not streaming.

Related Posts:







