Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Google Releases Gemini 3.8 TTS Models

Google Releases Gemini 3.8 TTS Models

Google DeepMind·Thursday, September 24, 2026
  • •Google released Gemini 3.8 Flash TTS and Flash-Lite TTS for speech generation on September 23, 2026.
  • •Flash TTS offers over 2,000 voices and voice replication from a 30-second sample with speaker consent verification.
  • •Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark; the models ranked first and second on its quality index.
  • •Google released Gemini 3.8 Flash TTS and Flash-Lite TTS for speech generation on September 23, 2026.
  • •Flash TTS offers over 2,000 voices and voice replication from a 30-second sample with speaker consent verification.
  • •Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark; the models ranked first and second on its quality index.
  • •Google released Gemini 3.8 Flash TTS and Flash-Lite TTS for speech generation on September 23, 2026.
  • •Flash TTS offers over 2,000 voices and voice replication from a 30-second sample with speaker consent verification.
  • •Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark; the models ranked first and second on its quality index.
  • •Google released Gemini 3.8 Flash TTS and Flash-Lite TTS for speech generation on September 23, 2026.
  • •Flash TTS offers over 2,000 voices and voice replication from a 30-second sample with speaker consent verification.
  • •Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark; the models ranked first and second on its quality index.

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, for developers, businesses and creators building generated speech experiences. The models are available through Google AI Studio and the Gemini API, with Flash TTS also in Gemini Notebook and Flash-Lite TTS in Google Vids. Enterprise API access is coming soon for both. Google said the models are its most expressive audio generation models yet and complement Gemini Audio products including 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking.

Flash TTS is designed for creative voice direction: users can build voices from scratch with natural-language prompts, choosing roles, accents and vocal traits across more than 100 languages and dialects. It provides over 2,000 production-ready voices, including Mexican Spanish, Quebec French and Scots English. Users can replicate a voice from a 30-second sample, provided they have the rights to use it; Google requires a verbal consent recording that matches the reference speaker. Custom voices can be saved for consistent use across projects. Voice remixing—adjusting timbre, pitch, pace and accent from an existing library voice—is marked as coming soon.

Both models let users direct dialogue line by line with cues for acting, pacing, dialect and nonverbal reactions such as laughter or sighs. Flash TTS supports long-form audio that maintains voice quality and character timbre across hours, as well as two-speaker scenes with distinct voices and conversational turn-taking. Flash-Lite is positioned for high-volume, cost-efficient dubbing, audio creation and expressive voice agents, with fine-grained control over tone, pacing and nuance.

Google said Flash TTS ranked #1 on Hume AI’s Voice Design Benchmark with a score of 71.4 and led accent modeling with 60.8. Flash TTS and Flash-Lite placed #1 and #2, respectively, on Hume AI’s Overall Quality Index. Google also reported improvements over Gemini 3.1 Flash TTS in long-form content and dual-speaker screenplay control. In blind preference evaluations on Voice Arena, the models took top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.

Google said every audio clip generated by Gemini Audio models carries an imperceptible SynthID watermark intended to make AI-generated speech detectable. Voice replication also uses consent verification. Developers can try the models in Google AI Studio; platforms including Agora, LiveKit, Pipecat and Vercel can use the Gemini API for speech-generation services. Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang are integrating the models for dubbing, media localization and voice agents. Voice replication in AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India.

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, for developers, businesses and creators building generated speech experiences. The models are available through Google AI Studio and the Gemini API, with Flash TTS also in Gemini Notebook and Flash-Lite TTS in Google Vids. Enterprise API access is coming soon for both. Google said the models are its most expressive audio generation models yet and complement Gemini Audio products including 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking.

Flash TTS is designed for creative voice direction: users can build voices from scratch with natural-language prompts, choosing roles, accents and vocal traits across more than 100 languages and dialects. It provides over 2,000 production-ready voices, including Mexican Spanish, Quebec French and Scots English. Users can replicate a voice from a 30-second sample, provided they have the rights to use it; Google requires a verbal consent recording that matches the reference speaker. Custom voices can be saved for consistent use across projects. Voice remixing—adjusting timbre, pitch, pace and accent from an existing library voice—is marked as coming soon.

Both models let users direct dialogue line by line with cues for acting, pacing, dialect and nonverbal reactions such as laughter or sighs. Flash TTS supports long-form audio that maintains voice quality and character timbre across hours, as well as two-speaker scenes with distinct voices and conversational turn-taking. Flash-Lite is positioned for high-volume, cost-efficient dubbing, audio creation and expressive voice agents, with fine-grained control over tone, pacing and nuance.

Google said Flash TTS ranked #1 on Hume AI’s Voice Design Benchmark with a score of 71.4 and led accent modeling with 60.8. Flash TTS and Flash-Lite placed #1 and #2, respectively, on Hume AI’s Overall Quality Index. Google also reported improvements over Gemini 3.1 Flash TTS in long-form content and dual-speaker screenplay control. In blind preference evaluations on Voice Arena, the models took top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.

Google said every audio clip generated by Gemini Audio models carries an imperceptible SynthID watermark intended to make AI-generated speech detectable. Voice replication also uses consent verification. Developers can try the models in Google AI Studio; platforms including Agora, LiveKit, Pipecat and Vercel can use the Gemini API for speech-generation services. Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang are integrating the models for dubbing, media localization and voice agents. Voice replication in AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India.

Read original (English)·Sep 23, 2026
#gemini 3 8 flash tts#flash lite tts#text to speech#voice generation#voice replication#synthid#hume ai benchmark#multilingual audio