Google Adds Talking Video Avatars to Gemini Enterprise Conversations
The avatars can speak across 97 languages and keep talking while fetching data. Businesses need allowlisted access to create their own.
Listen to this story
The audio brief
Story brief
3 key pointsGoogle’s September 24 enterprise rollout gives Gemini 3.8 Live a generated video persona that can take audio and visual input, speak across 97 languages, and keep a conversation going while background tool calls run. For companies, that opens a more embodied format for service and guided workflows—but a fluent avatar is not proof a lookup or action has completed. Preset faces are available, while reference-image...
- 01
Google describes video as near-real-time but provides no measured latency for this launch.
- 02
Custom reference-image avatars require enterprise allowlisting; preset avatars are also available.
- 03
SynthID is imperceptible in ordinary use, and the announcement leaves customer-facing disclosure choices unspecified.
An enterprise AI conversation can now have a face that moves and speaks as the exchange unfolds. Google introduced Gemini 3.8 Live with Live Avatar on September 24, pairing near-real-time generated video with speech and making it available in Gemini Enterprise. The pitch goes beyond a talking image: the avatar can take in audio and visual input while the conversation continues.
A voice model gets a visible persona
The addition follows last week’s Gemini 3.8 Live launch. Live Avatar gives that conversational model a video persona, with lip movements and expressions intended to track its speech. Google presents customer service and interactive walkthroughs as possible uses. In either setting, the change is what a person encounters during the exchange, not just what the AI says.
Google says the feature processes visual and audio input together, then responds with generated speech and video. That matters for a live exchange: an avatar meant to react to what someone shows it must coordinate what it says with what its face does. Google describes the video as near-real-time, rather than giving a measured response time for this launch.
The exchange continues during a data lookup
The avatar is designed to keep talking while it calls tools and fetches data in the background. Google calls this asynchronous tool execution: the system starts a separate task without waiting in silence for its result. Its example is a hotel guest check-in, where a lookup can run while the dialogue carries on.
A smooth conversation is not necessarily a finished task. If a lookup is still running, the avatar’s continued presence should not be mistaken for confirmation that the requested work is done. Google shows how the exchange is meant to continue, but the announcement does not provide a completion signal for users watching the avatar.
Language changes, and a gated custom face
Google says Live Avatar can switch among 97 languages during a conversation, adapting lip movements and expressions to match the speech. It also claims those switches do not degrade video fidelity or cause visual drift. These are Google’s descriptions of the feature, not independently measured results in the announcement.
Businesses can choose from preset avatars. Developers can also use a reference image to generate a responsive character that retains the image’s likeness or brand styling. That second route has a narrower entry point: custom avatar creation is available only through enterprise allowlisting. Google does not present it as a standard option for every Gemini Enterprise customer.
Making a generated face identifiable
Google says the generated audio and video carry SynthID, an imperceptible watermark intended to help identify AI-made content and limit misattribution. Because it cannot be seen or heard in an ordinary conversation, the watermark serves a different purpose from an obvious on-screen disclosure. Google points readers to a model card for its fuller safety approach.
That leaves a practical choice for businesses putting an avatar in front of people: how plainly to identify the character during the interaction. The launch establishes enterprise availability and describes the watermark, but it does not specify what a customer must tell someone speaking with the avatar. The visible design and the disclosure around it may shape the same encounter.
Sources
- blog.googleIntroducing Gemini 3.8 Live with Live Avatar
Reader comments
Newest comments first. Replies stay oldest first.