Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models in its Gemini Audio family. Google calls them its most expressive audio generation models yet. Flash TTS targets creative direction and character voices, while Flash-Lite TTS targets high-volume, cost-efficient production. Both let developers direct delivery line by line using natural language.
Is it deployable? Yes. Both models are rolling out now through the Gemini API and Google AI Studio. Access is API-only, with no open weights for self-hosting. Enterprise API access via Gemini Enterprise is listed as coming soon.
What Google Shipped
The release splits TTS into two tiers with shared direction controls:
- Gemini 3.8 Flash TTS is built for deep creative direction and character design. Target uses include gaming, immersive audiobooks, podcasts, and interactive media. It offers granular control over acting cues, pacing, dialect shifts, and backchanneling.
- Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale. Google positions it for dubbing, audio content creation, and expressive voice agents. It offers fine-grained control over tone, pacing, and expressive nuance.
In AI Studio, the playground links use the model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
Voice Design From a Text Prompt
Previous Gemini TTS offered 30 original voices. The 3.8 release moves to a much larger voice system.
- Generative voice design: Flash TTS creates new voices from prompts describing role, accent, and voice characteristics. This works across more than 100 languages and dialects. Goo...
via MarkTechPost
