RC RANDOM CHAOS

Google's Gemini 3.8 TTS adds prompt-built voices and 30-second voice cloning

· via Hacker News

Original source

Gemini 3.8 text-to-speech

Hacker News →

Google has released two new text-to-speech models under the Gemini 3.8 line. Flash TTS targets creative work — designing original character voices from natural-language prompts and directing delivery line by line — while Flash-Lite TTS is tuned for cheap, high-volume jobs like dubbing and voice agents. Both span 100-plus languages, ship a library of 2,000-plus prebuilt voices (up from 30), and support long-form narration, two-speaker dialogue, and nonverbal cues like laughs and interjections. Google cites top rankings on Hume AI’s voice-design and quality benchmarks against competitors.

The security-relevant feature is voice replication: the models can reconstruct a speaker’s voice from a 30-second sample. Google gates this behind a consent flow that requires a verbal authorization recording matching the reference speaker, and every generated clip carries an embedded SynthID watermark plus C2PA provenance credentials. Those safeguards are an implicit acknowledgment that convincing 30-second voice clones are now a commodity feature, with the obvious downstream risks for vishing, impersonation, and audio-based fraud.

The models are rolling out now through the Gemini API and Google AI Studio, with enterprise API access via Gemini Enterprise promised later. Consumer surfaces get them through Gemini Notebook and Google Vids, and platforms including LiveKit, Pipecat, Vercel, Figma, and HeyGen are already integrating them. How well the consent verification and watermarking hold up against determined abuse — rather than the audio quality — will be the story worth watching.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.