Google Releases Gemini Speech Models With Custom Voices and Line-by-Line Direction
Flash TTS can copy a permitted voice from a 30-second sample, but Google requires a matching consent recording. Enterprise API access is still to come.
Listen to this story
The audio brief
Story brief
3 key pointsGoogle’s September 23 launch gives developers two distinct speech-generation options: Flash TTS for designed voices and detailed performance direction, and Flash-Lite TTS for high-volume dubbing, audio production and voice agents. Flash can create reusable voices from descriptions or a 30-second recording, but copying requires a matching verbal consent recording; generated clips receive SynthID watermarks. Both...
- 01
Flash offers more than 2,000 ready-made voices and voice design across 100-plus languages and dialects; voice remixing is planned, not launched.
- 02
Both models support line-level cues, expressive sounds and two-speaker scenes; Google claims hours-long character consistency, which has not been independently tested.
- 03
Flash-Lite is positioned for production at scale, but Google provided no price comparison to substantiate a cost advantage.
A script for an AI voice no longer has to specify only the words. Google released two Gemini text-to-speech models on September 23 that let users direct how individual lines sound, including their pacing and emotion. One also creates voices from descriptions or copies a permitted voice from a recording, putting the consent check at the center of a more flexible production tool.
First choose a voice, then direct the scene
Gemini 3.8 Flash TTS handles the most detailed voice design. A user can describe a character’s role, accent and vocal qualities in ordinary language instead of selecting only from preset voices. Google says that design feature works across more than 100 languages and dialects. It also offers a library of more than 2,000 ready-made voices, including regional varieties such as Quebec French and Mexican Spanish.
Flash TTS can also make a reusable vocal profile from a 30-second sample of the user’s voice or one they have the right to use. Google says creators can save custom voices for later work. Voice remixing—adjusting a library voice’s pitch, pace, accent or other qualities—is planned, not part of the launch.
Both new models accept directions for individual lines. A writer can specify delivery through stage directions or script cues, then add laughter, sighs or brief listening sounds. A single script can stage a two-speaker exchange with separate voices and conversational turns. Google also says the models can sustain character voices across hours of audio with little drift, a useful claim for audiobooks and podcasts that has not been independently tested here.
Where the voices are rolling out
The two versions are rolling out to developers through the Gemini API and Google AI Studio. Studio’s new audio workspace lets developers prompt a voice, use voice replication and direct a two-speaker scene. Outside those developer tools, Google is putting Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids. API access through Gemini Enterprise is coming later.
Google describes Flash-Lite as the option for producing speech at scale, including dubbing and voice agents. It has not supplied a price comparison in this announcement. For teams choosing between the models, the stated split is therefore about intended workload, not a quantified cost advantage.
A consent check for voice copying
To replicate a voice, users must provide a verbal consent recording from the voice owner that matches the reference speaker, Google says. The company also says every clip generated by its Gemini Audio models receives an imperceptible SynthID watermark. It lists C2PA credentials among the protections for replicated voices. Those measures address two different points: permission to create a voice and a way to identify generated audio afterward.
The permission step matters because this is not limited to inventing fictional characters. Flash TTS can recreate a recognizable vocal profile from a short recording. Google describes the check, but its announcement does not measure how well it resists an attempt to misuse someone else’s voice.
What the quality claims show
Google says Flash TTS placed first on Hume AI’s Voice Design Benchmark, scoring 71.4 overall and 60.8 for accent modeling. It says Flash TTS and Flash-Lite TTS took the top two places on Hume AI’s Overall Quality Index. Those are company-reported results, rather than proof that either model will perform equally well on every script, speaker or language.
The more immediate test is the workflow Google has opened to developers: whether a designed voice stays consistent when a creator revises a long script, directs a conversation line by line and returns to that voice later. Google claims its tools can handle each part. The rollout now puts that claim in developers’ hands.
Sources
- blog.googleGemini 3.8 text-to-speech says hello
Reader comments
Newest comments first. Replies stay oldest first.