Google just made a strong case that voice AI is no longer a party trick. On September 15, 2026, the company announced two new models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, positioning them as production-ready infrastructure for voice agents and conversational AI. The timing matters. With OpenAI’s Realtime API, ElevenLabs, and a wave of voice-first startups competing for developer attention, Google needs more than benchmark wins. It needs models that actually work in the real world.
What each model does
The two models are built for different use cases. Gemini 3.8 Live is the cost-efficient option, optimized for scale and fluid dialogue. It processes visual inputs in near real-time, supports automatic language detection and switching across 97 languages mid-conversation, and can execute tool calls and API requests in the background while keeping the conversation going. That last feature is more useful than it sounds. Most voice models either pause or drop context when running background tasks. This one doesn’t.
Gemini 3.8 Live Extended Thinking is built for complexity. It reasons and speaks at the same time, using verbal cues like ‘Let me check that…’ to bridge the gap between thinking and responding. It also narrates progress through multi-step tasks live, which reduces the awkward silences that plague most voice agents today. Think of it as the model you’d deploy when a user is asking an AI to coordinate a multi-step booking, debug a codebase by voice, or build a business plan on the fly.
How the benchmarks stack up
Google is leaning hard into the numbers here. Gemini 3.8 Live Extended Thinking claims the top spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It scores 68.6% on τ-Voice for agentic task completion and 35.1% on Sierra’s τ-Voice-banking benchmark. On Big Bench Audio, it hits 97.7%. These are strong figures, and they put it ahead of current alternatives in most categories.
Gemini 3.8 Live, the cheaper sibling, placed second in the Speech Agent Arena and holds its own on ServiceNow’s EVA-Bench, which measures voice agents across complex workflows. Second place in an arena benchmark might sound like a consolation prize, but for a model priced for scale and enterprise deployment, it’s a solid position to be in.
Developer and enterprise access
Google is making both models available through several channels starting today:
- Gemini 3.8 Live: Available in the Gemini API and Google AI Studio for developers, in private preview in Gemini Enterprise, and in Search Live for general users
- Gemini 3.8 Live Extended Thinking: Available in the Gemini API and Google AI Studio, in private preview in Gemini Enterprise, and in Gemini Live for all users, with Docs access for Google AI Pro and Ultra subscribers, and Gmail and Keep access for all Google AI subscribers
On the developer infrastructure side, Google is working with Agora, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents to let teams build on top of the Live API without managing real-time media streaming themselves. Enterprise partnerships with Salesforce, Genspark, and Lumeris are also in play, with those companies citing latency and tool-calling as key reasons for interest.
Safety and what it means for the market
All audio output is watermarked using Google’s SynthID technology. The watermark is embedded directly in the audio and is imperceptible to listeners, but detectable by systems designed to identify AI-generated content. That’s a meaningful feature as synthetic voice content becomes harder to distinguish from human speech.
So where does this leave the field? OpenAI’s Realtime API is still the default choice for many developers, and ElevenLabs owns a strong position in voice quality for non-agentic use cases. But Google now has something neither competitor has fully demonstrated: a model that can reason, speak, see, and run background tasks simultaneously without losing conversational thread. If that holds up in production, it’s a real differentiator. The benchmark lead is notable. The real test is whether developers build on it.



