Google just made every static voice library feel outdated. The company announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models that go far beyond picking from a dropdown of preset voices. You can now describe a voice in plain language, and the model builds it. A fire-breathing dragon with a gravelly baritone, a high-energy Melbourne DJ, a monotone robot. These are not metaphors. Those are the actual demo examples Google published.
What the two models actually do
The split between the two models follows a familiar Google pattern: one for quality and creative depth, one for scale and cost efficiency.
Gemini 3.8 Flash TTS is the flagship. It’s built for creators, developers, and anyone who needs precise control over how a voice performs. You can write stage directions into the script, specify pacing and dialect, add non-verbal cues like laughter or backchanneling interjections, and direct two-speaker conversations from a single script. It also supports voice replication from a 30-second audio sample, with consent verification baked in.
Gemini 3.8 Flash-Lite TTS is aimed at high-volume use cases like dubbing, voice agents, and content pipelines where cost per character matters. It still supports expressive control over tone and pacing, but the priority is throughput, not fine-grained creative direction.
Both models support over 100 languages and dialects, including regional varieties like Mexican Spanish, Quebec French, and Scots English. The voice library has over 2,000 production-ready options, which is a significant step up from the 30 voices Google previously offered.
How it compares to the competition
ElevenLabs has been the go-to for voice cloning and expressive TTS. OpenAI’s voice API is solid but relatively limited on customization. Cartesia and PlayHT are popular in the developer community for low-latency use cases. Google is now coming at all of them with a broader feature set and scale advantages from its existing API infrastructure.
On Hume AI’s Voice Design Benchmark, Gemini 3.8 Flash TTS scored 71.4 overall, taking the top spot. It also led specifically in accent modeling with a score of 60.8. Both models took the top two positions on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, the models ranked highly in Japanese, Brazilian Portuguese, Vietnamese, Arabic, Mexican Spanish, and Hindi. Those are not easy languages to score well in for TTS, and regional accent accuracy is where most competitors still struggle.
Key features worth knowing
- Generative voice design: describe a character and the model creates the voice using natural language prompts
- Voice replication from a 30-second sample, with verbal consent verification and SynthID watermarking
- 2,000-plus production-ready voices with broad regional language coverage
- Line-by-line performance direction with support for vocal bursts, laughter, and active-listening interjections
- Native two-speaker scene staging from a single script
- Long-form generation with minimal speaker drift, suitable for full audiobooks or podcast episodes
- Voice remixing is coming soon, letting users adjust timbre, pitch, and accent on existing library voices
Where you can use it
Gemini 3.8 Flash TTS is available now in the Gemini API and Google AI Studio. Enterprise users get access through Gemini Enterprise, Gemini Notebook, and Google Vids. Flash-Lite TTS is rolling out on a similar timeline.
Google is also working with integration partners including Figma, HeyGen, Agora, LiveKit, Pipecat, and Vercel. Those partnerships matter because they signal real developer adoption pathways, not just API documentation sitting in a vacuum.
Why this matters
Voice is becoming infrastructure. Conversational AI agents need it. Dubbing pipelines need it. Audiobook and podcast production is increasingly automated. The companies that own the best voice generation layer will have serious leverage as these workflows scale.
Google’s move here is not just about adding features. It’s about positioning Gemini Audio as a full stack, from live translation and transcription to now generative TTS with performance direction. The SynthID watermarking on every output is also a meaningful trust signal for enterprise buyers who are increasingly worried about synthetic voice misuse.
For developers evaluating TTS options right now, Gemini 3.8 Flash TTS is worth a serious look, especially if multilingual accuracy or creative voice control are on your requirements list.




