Google has released two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, bundling more than 2,000 prebuilt voices and a custom voice feature that requires only a 30-second audio sample. Simon Willison's Gemini TTS Playground lets you test both models immediately in-browser, and Google's official announcement confirms the models are available through the Gemini API today.
Why it matters
Text-to-speech has historically lived in a separate API silo, meaning teams building voice AI features had to stitch together an LLM call and a TTS call from different providers, managing latency, auth, and billing across two surfaces. Folding TTS directly into the Gemini model family collapses that stack. The 2,000-voice library also moves the default well past the handful of preset voices that have been the norm since ElevenLabs popularized the category back in 2023.
Custom voice from 30 seconds of audio is the detail that changes the product conversation, not the voice count.
The lite variant follows the pattern Google has used across the Gemini 3.x line: a cheaper, faster option for latency-sensitive or high-volume use cases where top-tier fidelity is not required.
What changes in practice
- Teams already on the Gemini API can add voice output without a new vendor contract or SDK.
- Custom voice personas for agents, assistants, or branded products now have a much lower data requirement (30 seconds vs. minutes of studio audio).
- The two-tier model structure (flash vs. flash-lite) means cost optimization is a first-class concern from day one, not an afterthought.
- Multimodal prompt engineering now has a voice output dimension to design around, including tone, pacing, and persona consistency across turns.
How to use it
- Prototype in the playground first. Use Willison's tool to audition voices and test your actual script copy before writing any integration code.
- Record your custom voice sample carefully. 30 seconds is a tight window. Use a clean, noise-free recording with natural sentence variety. Monotone samples produce flat output.
- Start with flash-lite for non-critical paths. Notifications, confirmations, and low-stakes agent responses are good candidates. Reserve flash for conversational or high-visibility surfaces.
- Design prompts for spoken output, not reading. Sentences that work on screen often sound unnatural when spoken. Rewrite for rhythm: shorter sentences, no parentheticals, no markdown symbols in the text you pass to the model.
- Benchmark latency end-to-end, not just TTS latency. If your pipeline chains an LLM call into a TTS call, measure the full round-trip. Streaming output from the LLM directly into the TTS call can cut perceived latency significantly.
The consolidation of LLM and TTS into one API is a workflow change as much as a capability change. Teams that adapt their prompt and pipeline design now will have a head start when Google inevitably extends this pattern to speech input as well.
READY TO ASCEND
Get AI news that respects your time
The signal, distilled. Curated AI news and prompt-engineering insight. No noise.