Intelligent transcription with Gemini 3.5 Transcribe

Today marks a significant advancement in the realm of voice technology with the unveiling of Gemini 3.5 Transcribe, a state-of-the-art speech-to-text model meticulously crafted for intelligent voice interactions. This innovative model transcends the limitations of traditional speech recognition systems, which often falter in the presence of background noise, intricate jargon, and the need for disfluency cleanup. Gemini 3.5 Transcribe is adept at transforming raw audio into precise, polished, and formatted text.

Users of the Gemini app and Android devices have already begun to experience the advantages of this cutting-edge transcription model, particularly through new voice functionalities such as Rambler on Android and within the Gemini app on macOS. Developers are now empowered to harness similar capabilities via the Gemini 3.5 Transcribe in the Gemini API, available in Google AI Studio and the Gemini Enterprise Agent Platform.

Gemini 3.5 Transcribe has been engineered for seamless integration into developer workflows, catering to a variety of applications including voice agents, real-time captioning tools, and post-call analytics pipelines. The model is accessible through two distinct APIs:

  • Real-time streaming: Offers continuous, bidirectional streaming with sub-second latency for interactive voice applications via the Live API using gemini-3.5-transcribe-live.
  • Pre-recorded audio processing: Facilitates the transcription of recorded audio, meetings, call logs, and more, complete with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

Get more precise and intelligent transcription

Gemini 3.5 Transcribe is meticulously designed to capture the nuances of natural speech, enhancing its ability to comprehend intent and recognize specialized vocabulary, thus enabling users to perform tasks vocally with greater ease.

  • Smart transcription: Effectively manages self-corrections (for instance, “let’s meet Tuesday—no, Wednesday”), eliminates filler words like “ums” and “ahs,” and auto-formats the resulting text.
  • Function calling: The model can assign complex tasks, such as image generation and file analysis, to other Gemini models through function calls, currently available in the Gemini macOS app.
  • More precise transcription: According to Artificial Analysis, it boasts an impressive average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming scenarios, demonstrating robust performance even in noisy, real-world settings while accurately capturing alphanumeric entities like postal codes and order IDs.
  • Custom vocabulary: Adapts to specialized jargon and unique spellings, ensuring that transcriptions align with the provided custom vocabulary.
  • Global language support: Automatically identifies and transcribes over 85 languages, adeptly managing regional accents and diverse dialects.
  • Multi-speaker identification: Accurately attributes speech in pre-recorded audio, providing timestamps for up to three speakers, with experimental support for more than three speakers.
AppWizard
Intelligent transcription with Gemini 3.5 Transcribe