Gemini 3.5 Transcribe is a new speech-to-text model designed for intelligent voice interactions, overcoming challenges faced by traditional speech recognition systems, such as background noise and complex jargon. It transforms raw audio into accurate, formatted text and is available to users of the Gemini app and Android devices. Developers can access its capabilities through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
The model supports two APIs: real-time streaming for interactive applications with sub-second latency and pre-recorded audio processing that includes speaker attribution and word-level timestamps. It enhances transcription accuracy by managing self-corrections, eliminating filler words, and auto-formatting text. The model has a Word Error Rate of 4.0% for streaming and 2.6% for non-streaming scenarios, effectively capturing alphanumeric entities. It adapts to custom vocabulary, supports over 85 languages, and can identify multiple speakers in pre-recorded audio.