Most AI voice agents sound fine but feel hollow. There’s nobody home. Google’s answer to that problem is Live Avatar, a feature built into Gemini 3.8 Live that pairs a speaking, lip-synced video persona with the model’s real-time dialogue capabilities. As Google announced on September 24, 2026, this is now available through Gemini Enterprise — and it’s a direct push into territory that startups like Synthesia, HeyGen, and D-ID have been building toward for years.
What Live Avatar actually does
Live Avatar generates a video persona that listens, responds, and reacts with synchronized lip movements and facial expressions during a live conversation. It’s not a pre-rendered clip. The avatar is generated in near real time alongside the audio output, creating something closer to a video call than a scripted demo.
The feature is multimodal in both directions. It processes audio and visual input simultaneously, so the avatar isn’t just responding to text. It’s reacting to what it hears and sees. Google says the system supports natural turn-taking, which matters more than it sounds — choppy interruptions are one of the fastest ways to break the illusion of presence in any conversational interface.
The technical details worth paying attention to
A few specific capabilities stand out here:
- Asynchronous tool calling: the avatar can trigger background tasks and fetch data while continuing the conversation without a visible pause
- Multilingual support across 97 languages with real-time lip-sync adaptation and expression adjustment per language
- Custom avatar creation from a single high-quality reference image, preserving likeness and brand styling
- SynthID watermarking embedded in both the audio and video output for AI content detection
The asynchronous tool calling piece is genuinely useful for enterprise applications. Think of a hotel check-in agent that can query a reservation system in the background without going quiet for three seconds while it waits. That kind of responsiveness is what separates a working product from a demo.
How it fits into the competitive picture
Google is entering a space that already has real players. Synthesia and HeyGen have built businesses on AI-generated video presenters, and ElevenLabs has been pushing hard on expressive voice. But those tools are mostly async. You render a video, you export it, done. Live Avatar is built for real-time interaction, which is a harder engineering problem and a different market.
The closest comparable is probably Tavus, which has been working on real-time conversational video personas for enterprise use. But Google has something Tavus doesn’t: a full-stack AI platform underneath it. Pairing Live Avatar with Gemini’s reasoning, tool use, and native multilingual models is a more integrated bet than stitching together third-party APIs.
Availability and what to expect
Live Avatar is currently available in Gemini Enterprise. Custom avatar creation requires enterprise allowlisting, so it’s not a self-serve feature yet. That’s probably deliberate. Identity-based generation carries real risk, and Google’s SynthID watermarking approach suggests they’re thinking carefully about misuse.
For developers evaluating this, the honest question is whether enterprise customers will actually trust a video avatar in production-facing roles. Customer service is the obvious entry point. But the 97-language support hints at something bigger: global-scale deployments where a consistent visual presence across markets would otherwise require significant localization work. That’s a real cost problem for multinationals, and this could genuinely address it.




