For live applications, developers can use gemini-3.5-transcribe-live through the Live API. Google says that version supports continuous bidirectional streaming with latency below one second.
Recorded audio is handled through the Interactions API using gemini-3.5-transcribe. That version is intended for material such as meetings and call logs and includes speaker identification and word-level timestamps.
According to measurements from Artificial Analysis cited by Google, Gemini 3.5 Transcribe records an average word error rate of 4.0% for streaming use and 2.6% for non-streaming transcription. Google also says the time required to produce a final transcription is 70% lower than with its previous Chirp 3 model.
Performance remains different across multilingual testing. On the FLEURS benchmark, Google reports a 5.50% word error rate for streaming and 5.04% for non-streaming use across a selection of languages and locales, improving on Chirp 3.
Language coverage extends to more than 85 languages, with automatic detection and support for regional accents and dialects. For recorded audio, the model can identify as many as three speakers, while attribution beyond three speakers remains experimental.
Google is also positioning the model around more natural dictation. Gemini 3.5 Transcribe can interpret corrections made while someone is speaking, such as changing a date mid-sentence, and remove verbal fillers before producing the final text. The model can also recognize alphanumeric information such as postal codes and order IDs in noisy audio, according to Google. Developers can provide their own vocabulary lists when applications need more reliable handling of domain-specific terminology.
Function calling adds another layer beyond transcription. Gemini 3.5 Transcribe can hand off tasks such as file analysis or image generation to other Gemini models. Google says that capability is currently available in the Gemini app on macOS. The same model is already being used across several Google products. Rambler on Gboard for Android uses Gemini 3.5 Transcribe to turn spoken input into formatted text and lets users make edits or adjust writing style using voice commands.
In Google Antigravity, the transcription system can use screen context and chat history, with user permission, to improve recognition of information such as file names and active document content.
Google AI Studio is also using the model in Build mode, where developers can create applications through spoken instructions. In the Gemini app on macOS, users can combine voice input with screen context to work with local files, move text between applications or generate images.
Chrome is next on Google’s roadmap. The company says voice typing will eventually be available in web fields, allowing users to dictate replies, posts and prompts directly in the browser.
Google is also working with developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, which can use the Gemini Live API to provide the media-streaming infrastructure behind voice applications. Companies including Vivo, Intellitek Health and Lingopal have also been testing the model, according to Google.
With Gemini 3.5 Transcribe, Google is tying speech recognition more closely to the rest of the Gemini ecosystem. The model combines lower-latency transcription with formatting, contextual understanding, multilingual support and function calling, giving developers a single speech layer that can feed directly into broader voice-driven workflows.
This analysis is based on reporting from Google.
Image courtesy of Google.
This article was generated with AI assistance and reviewed for accuracy and quality.