Voice Synthesis by Description
Google has officially rolled out two text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, spanning over 100 languages. According to the release documentation, Gemini 3.8 Flash TTS is aimed squarely at creative production—podcasts, audiobooks, and game development—by allowing operators to spin up entirely unique vocal identities directly from text prompts. Users can still pull from a catalog of over 2,000 preset voices featuring regional accents like Mexican Spanish, Quebec French, and Scottish English, but the prompt-based generation is where the real leverage lies.
Beyond pure generation, the platform bundles a voice cloning utility that builds a functional profile off a brief 30-second audio sample. Google also introduced Voice Remixing for tweaking timbre, pitch, tempo, and accent on existing library assets. To keep the content tracking departments happy, every clip baked by these systems ships with an inaudible SynthID watermark.
Script Control and Dialogue Mechanics
Both models parse fine-grained script directions, letting engineers inject nonverbal audio cues like sighs, laughter, and conversational filler such as "mhm" straight into the text stream. The underlying architecture handles two-voice scripts natively within a single generation pass, preventing the awkward character bleed that usually ruins long-form audio generation.
"Flash TTS is aimed at creative uses such as podcasts, audiobooks, and game characters, while Flash-Lite TTS is designed for low-cost speech generation at scale for dubbing, audio content, and voice agents."
This setup gives automated contact centers and dubbing pipelines a way to push out high-volume conversational audio without bleeding budget on human talent. Flash-Lite TTS in particular is explicitly optimized for high-throughput, low-latency deployments where cost-per-minute dictates the unit economics.
Generation Economics and Platform Availability
Google is making both models accessible through the Gemini API and Google AI Studio. For businesses running high-volume customer service scripts or localization pipelines, this release signals a hard ceiling on what human voice actors can charge for commodity reads.