Google Releases Gemini Speech Models With Custom Voices and Line-by-Line Direction

Flash TTS can copy a permitted voice from a 30-second sample, but Google requires a matching consent recording. Enterprise API access is still to come.

By 3 min read
Google Releases Gemini Speech Models With Custom Voices and Line-by-Line Direction
Google Releases Gemini Speech Models With Custom Voices and Line-by-Line Direction

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
Google’s new speech models let developers direct not just what a voice says, but how each line is delivered—and one can recreate a voice from a 30-second recording, provided its owner supplies a matching consent recording. The launch comes in two versions. Gemini 3.8 Flash TTS is the more expressive option: creators can describe a voice in ordinary language, choose from more than two thousand ready-made voices, and add line-by-line cues for pacing, emotion, or sounds such as laughter and sighs. Flash-Lite TTS is aimed at high-volume work like dubbing and voice agents. Both models can stage two-speaker conversations. Flash can also save a reusable voice profile from a short sample. That makes the consent requirement especially important: Google says the verbal recording must match the person in the reference audio. It says generated clips carry an imperceptible SynthID watermark, and lists C2PA credentials among protections for replicated voices. Those are safeguards Google describes; the announcement does not test how resistant the process is to misuse. The models are rolling out through the Gemini API and Google AI Studio. Flash TTS is also rolling out in Gemini Notebook, and Flash-Lite in Google Vids. Enterprise API access is still pending. Google claims voices can stay consistent over hours, but that has not been independently tested. It also gives no price comparison to support a cost advantage for Lite. The key test now is whether a designed voice holds together across a long, revised production.

Story brief

3 key points

Google’s September 23 launch gives developers two distinct speech-generation options: Flash TTS for designed voices and detailed performance direction, and Flash-Lite TTS for high-volume dubbing, audio production and voice agents. Flash can create reusable voices from descriptions or a 30-second recording, but copying requires a matching verbal consent recording; generated clips receive SynthID watermarks. Both...

  1. 01

    Flash offers more than 2,000 ready-made voices and voice design across 100-plus languages and dialects; voice remixing is planned, not launched.

  2. 02

    Both models support line-level cues, expressive sounds and two-speaker scenes; Google claims hours-long character consistency, which has not been independently tested.

  3. 03

    Flash-Lite is positioned for production at scale, but Google provided no price comparison to substantiate a cost advantage.

A script for an AI voice no longer has to specify only the words. Google released two Gemini text-to-speech models on September 23 that let users direct how individual lines sound, including their pacing and emotion. One also creates voices from descriptions or copies a permitted voice from a recording, putting the consent check at the center of a more flexible production tool.

First choose a voice, then direct the scene

Gemini 3.8 Flash TTS handles the most detailed voice design. A user can describe a character’s role, accent and vocal qualities in ordinary language instead of selecting only from preset voices. Google says that design feature works across more than 100 languages and dialects. It also offers a library of more than 2,000 ready-made voices, including regional varieties such as Quebec French and Mexican Spanish.

Flash TTS can also make a reusable vocal profile from a 30-second sample of the user’s voice or one they have the right to use. Google says creators can save custom voices for later work. Voice remixing—adjusting a library voice’s pitch, pace, accent or other qualities—is planned, not part of the launch.

Both new models accept directions for individual lines. A writer can specify delivery through stage directions or script cues, then add laughter, sighs or brief listening sounds. A single script can stage a two-speaker exchange with separate voices and conversational turns. Google also says the models can sustain character voices across hours of audio with little drift, a useful claim for audiobooks and podcasts that has not been independently tested here.

Where the voices are rolling out

The two versions are rolling out to developers through the Gemini API and Google AI Studio. Studio’s new audio workspace lets developers prompt a voice, use voice replication and direct a two-speaker scene. Outside those developer tools, Google is putting Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids. API access through Gemini Enterprise is coming later.

Google describes Flash-Lite as the option for producing speech at scale, including dubbing and voice agents. It has not supplied a price comparison in this announcement. For teams choosing between the models, the stated split is therefore about intended workload, not a quantified cost advantage.

A consent check for voice copying

To replicate a voice, users must provide a verbal consent recording from the voice owner that matches the reference speaker, Google says. The company also says every clip generated by its Gemini Audio models receives an imperceptible SynthID watermark. It lists C2PA credentials among the protections for replicated voices. Those measures address two different points: permission to create a voice and a way to identify generated audio afterward.

The permission step matters because this is not limited to inventing fictional characters. Flash TTS can recreate a recognizable vocal profile from a short recording. Google describes the check, but its announcement does not measure how well it resists an attempt to misuse someone else’s voice.

What the quality claims show

Google says Flash TTS placed first on Hume AI’s Voice Design Benchmark, scoring 71.4 overall and 60.8 for accent modeling. It says Flash TTS and Flash-Lite TTS took the top two places on Hume AI’s Overall Quality Index. Those are company-reported results, rather than proof that either model will perform equally well on every script, speaker or language.

The more immediate test is the workflow Google has opened to developers: whether a designed voice stays consistent when a creator revises a long script, directs a conversation line by line and returns to that voice later. Google claims its tools can handle each part. The rollout now puts that claim in developers’ hands.

Sources

  1. blog.googleGemini 3.8 text-to-speech says hello

Loading discussion...

YOUR READING SPACE

Notifications