Google’s Gemini Audio team shipped a pair of speech-generation models on September 23, replacing a fixed library of preset voices with a system that builds a voice from a written description. The bigger of the two, Gemini 3.8 Flash TTS, targets creative direction: a director can type an accent, a pacing style, or an emotional register and then adjust each line of a script on its own. Gemini 3.8 Flash-Lite trades that granularity for throughput. It is built for dubbing a large video catalog or running a voice agent at scale, jobs where volume matters more than nuance.

One feature inside Flash TTS carries the most risk: voice replication. The system can rebuild a vocal profile from a 30-second clip of “your voice or a voice you have the rights to use,” reusing it on demand. Google says the feature sits behind a consent step: before it will build the profile, the person whose voice is being cloned must supply a spoken recording that lines up with the reference clip. Every clip either model outputs also carries a SynthID watermark and C2PA content credentials, Google’s standard pairing for flagging synthetic media.

That consent step is a claim about intent more than a tested guarantee. Google’s post describes what the system requires, a matching spoken recording, but does not explain how the match is scored, what happens if someone reuses an old consent clip, or whether the check can be beaten by another synthetic voice reading the same phrase. Teams planning to build voice-cloning features on top of Flash TTS should ask Google for the verification method before shipping anything that touches a real person’s likeness.

On quality, every number in the launch post is Google’s own. The company says Flash TTS scores highest overall on Hume AI’s benchmark for voice design, at 71.4, and leads accent modeling separately at 60.8. Flash TTS and Flash-Lite also take first and second place, respectively, on Hume’s Overall Quality Index. Both models top blind human preference tests on Voice Arena in several languages, among them Japanese, Brazilian Portuguese, and Hindi. Hume AI and Voice Arena are independent evaluators, but Google chose which results to publish and did not release a full leaderboard alongside the post.

The release extends a Gemini Audio lineup that already spans Live Translate, Transcribe, and Live Extended Thinking, and Google is pushing distribution hard on day one. Flash TTS ships today in the Gemini API, Google AI Studio, and Gemini Notebook; access through Gemini Enterprise is listed as coming soon. Flash-Lite follows the same rollout path into Google Vids. Agora, LiveKit, Pipecat, and Vercel have already built support into their developer platforms. Google names Figma and HeyGen among its launch partners, alongside dubbing specialists Wondercraft and Linguana. The announcement gives no pricing for either model.

A voice-remixing tool is coming later: pick an existing library voice and nudge its accent or pace with a prompt rather than designing one from scratch. For any team building dubbing, audiobook, or voice-agent products, the number worth pressing Google on is not the benchmark score. It is the failure rate of that consent check against an adversarial sample, since that gate decides whether Flash TTS is safe to put in front of paying customers.

Google, “Gemini 3.8 text-to-speech says hello,” a company blog post from the Gemini Audio team, September 23, 2026.