Skip to content

Google Ships Two New TTS Models Built for Production Voice Agents

Short answer

Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026, offering custom voice creation, 2,000+ prebuilt voices, and line-by-line delivery control via the Gemini API. Enterprise API access through Gemini Enterprise is coming soon, meaning developers can build voice agent integrations now but enterprise procurement paths aren't fully open yet.

What this means for operators

For a support or sales team considering a voice agent, this changes what's technically possible today: the Gemini API is live now for developers to build custom-branded voices, multilingual dubbing, or conversational agents with realistic pacing and backchanneling, and Flash-Lite is explicitly pitched for high-volume, cost-efficient voice agents. The catch is that Gemini Enterprise access is still "coming soon" via API, so a 10-200 person company without in-house developers will likely need to wait or work through a partner platform (Agora, LiveKit, Pipecat, Vercel are named integrators) rather than get this through an enterprise console today.

Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, expanding its Gemini Audio family with what it calls its most expressive voice generation models to date.

Flash TTS is built for creative direction: developers can generate entirely new character voices from natural language prompts, direct delivery line by line (pacing, dialect, acting cues), and script two-speaker scenes with vocal bursts like <laughs> or <sigh> for realistic turn-taking. Flash-Lite TTS is positioned for high-volume, cost-efficient use cases such as dubbing and voice agents, with the same fine-grained tone and pacing controls.

Both models support over 100 languages and dialects, and Google says the library scales from 30 original voices to more than 2,000 production-ready voices, including regional varieties like Mexican Spanish, Quebec French, and Scots English. Voice replication lets a user recreate a consistent vocal profile from a 30-second sample, gated by consent verification, SynthID watermarking, and C2PA credentials. Voice remixing — fine-tuning an existing library voice with prompts — is coming soon. Voice replication is not available in Illinois, Texas, the EEA, UK, Switzerland, or India.

On benchmarks, Google reports Flash TTS took the top spot on Hume AI's Voice Design Benchmark (71.4) and led accent modeling (60.8), while Flash and Flash-Lite placed first and second respectively on Hume AI's Overall Quality Index. In blind evaluations on Voice Arena, both models led among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Rollout starts today for developers through the Gemini API and Google AI Studio for both models. Enterprise access via API in Gemini Enterprise is listed as coming soon. Flash TTS also reaches general users through Gemini Notebook, while Flash-Lite TTS reaches general users through Google Vids. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are enabling deployment of these models, and Google names Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang as partners integrating the TTS models for dubbing, localization, and conversational voice agents.

Source: Google DeepMind

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.