Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Google’s new Flash TTS models let you design AI voices from scratch using text descriptions

AI News

Google’s new Flash TTS models let you design AI voices from scratch using text descriptions

Google’s new Flash TTS models let you design AI voices from scratch…

Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue.

Gemini 3.8 Flash TTS is designed for creative projects such as game characters, audiobooks, and podcasts, while Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents, according to Google. Flash TTS lets users create voices from text descriptions With Gemini 3.8 Flash TTS, users can design voices from scratch. According to Google, a text prompt can define a voice’s role, accent, and vocal traits across a wide range of languages and dialects. For users who don’t want to start from zero, Google offers a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English.Ad A voice cloning feature can build a voice profile from a 30-second audio sample. To use it, the person whose voice is being cloned has to record a spoken statement of consent, and the voice in that recording must match the sample. Every clip the Gemini audio models generate carries an inaudible SynthID watermark to help detect AI-generated speech, according to Google.Ad Google has also announced “Voice Remixing,” a feature that will let users adjust the timbre, pitch, tempo, and accent of library voices, but it isn’t available yet. Script directions give users control over dialogue and delivery Both models let users write directions for each line or have the model interpret script cues on its own. Google says the models can generate hours of audio with minimal “speaker drift,” meaning the voice barely changes over time.Ad A two-voice mode generates dialogue from a single script while keeping the voices distinct, according to the company. Users can also script laughter, sighs, and sounds like “mhm” to place reactions and pauses exactly where they want them. In two of my own tests, I used a preset voice with a style prompt asking it to imitate an annoyed Berliner speaking English with a thick German accent. The style controls produced a convincing accent and intonation, but both tests had a high-pitched whine in the background at some points. In one of them, the voice also changed at the end of the clip.Ad Style prompt: A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. “Th” becomes “z” or “d” (“ze,” “sink,” “dat”), “w” becomes “v” (“vat,” “vell”), and final consonants become harder (“goot,” “bat” for “bad”). The “r” is guttural and throaty, never the English “r.” Vowels are flat and short, lacking the softness found in American or British English.