Imagine recording a speech or interview without having to worry about speech hesitations showing up in your notes. Google has introduced Gemini 3.5 Transcribe, an artificial intelligence model designed to make voice transcriptions cleaner and more professional by automatically filtering out filler words like "um" and "uh."
The new model represents a major advancement over Google's previous transcription tool, Chirp 3, particularly in multilingual performance and reducing wording error rates. In addition to stripping away unwanted verbal slips, Gemini 3.5 Transcribe automatically formats spoken text into readable paragraphs. It supports more than 85 languages and can identify up to three distinct speakers in pre-recorded audio, complete with word-level timestamps.
To prevent crucial terminology from being altered, the system allows users to upload custom vocabularies. This means specialized industry jargon, technical terms, and unique names are accurately transcribed without requiring tedious manual corrections afterward.
The feature is initially rolling out in English for macOS users on the Gemini app and on Android through the Rambler dictation feature in select regions. Developers can also access the tool in public preview via the Gemini API, Google AI Studio, and Antigravity, with web support for Chrome coming soon.