Google says Extended Thinking can reason and speak at the same time instead of forcing a conversation to stop while the model works through a task. It can acknowledge a request with phrases such as “Let me check that…” while continuing to process steps, call tools and make API requests in the background.
“Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet,” said Tom Ouyang, principal engineer at Google DeepMind.
The base Gemini 3.8 Live model also adds visual understanding during conversations. It can process visual inputs in near real time and automatically switch among 97 supported languages during the same conversation.
Google is backing the release with a series of benchmark results. Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index, taking the top overall position in that evaluation. It also recorded 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark for agentic task completion.
The company reported a 97.7% score on Big Bench Audio and said Gemini 3.8 Live placed second in the Speech Agent Arena, which measures user preferences for conversational systems.
Google also tested the models on ServiceNow’s EVA-Bench through the Gemini Enterprise Agent Platform. The company said the results showed a stronger balance between task accuracy and conversational quality on complex workflows.
“These models handle complex reasoning, real-time visual context, and background task execution without interrupting your conversation,” said Malini Jaganathan, a member of the Gemini Audio Team.
Google demonstrated several examples of that approach. One showed Gemini 3.8 Live using visual context while playing chess. Another used Extended Thinking to turn sketches and voice instructions into working React components. The company also showed the models coordinating multi-step reservations and generating business plans through spoken interaction.
For developers, Google is expanding the ecosystem around the Gemini Live API. Platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents can handle the underlying real-time media infrastructure while developers build voice-driven applications on top.
Salesforce, Genspark and Lumeris are among the companies Google identified as partners testing the models, with the company highlighting latency, conversational flow and tool-calling as key capabilities.
Google is also carrying the models into its own consumer and workplace products. Gemini 3.8 Live is rolling out to Search Live, while Extended Thinking is coming to Gemini Live and Google Workspace experiences including Docs, Gmail and Keep. Google AI Pro and Ultra subscribers can access Extended Thinking in Docs, while Google AI subscribers can use it in Gmail and Keep.
Enterprise access is beginning through private preview in Gemini Enterprise. Google also plans to bring the models to Gemini Enterprise for Customer Experience, with Extended Thinking additionally coming to Google Workspace business customers.
The company says pricing is part of the split between the two models. Gemini 3.8 Live is intended to offer lower-cost deployment at scale, while Extended Thinking targets heavier reasoning workloads at what Google describes as a competitive price compared with other frontier voice models.
All audio generated by Google’s AI products using these models carries SynthID watermarking. The watermark is embedded into the audio so AI-generated output can be detected later.
The release gives Google a two-tier voice stack built around the same idea: keep a conversation moving while the model works. Gemini 3.8 Live focuses on speed and scale, while Extended Thinking adds deeper reasoning for workflows that require multiple steps, background actions and more sustained problem-solving.
This analysis is based on reporting from Google.
Image courtesy of Google.
This article was generated with AI assistance and reviewed for accuracy and quality.