Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/A new generation of voice gemini 38 flash tts for voice ai developers 2
Dev48

© 2026 · All rights reserved.

A New Generation of Voice: Gemini 3.8 Flash TTS for Voice AI Developers

Источник: Agora

A New Generation of Voice: Gemini 3.8 Flash TTS for Voice AI Developers

Source: Agora

Explore Gemini 3.8 text-to-speech, expressive voice controls, and what the latest speech generation capabilities could mean for developers building voice AI.

September 28, 2026•Updated: September 28, 2026

A voice can change the feel of an entire conversation. The same words can sound reassuring, excited, uncertain, or sincere depending on how they are delivered. For developers building voice experiences, getting that delivery right has often meant working around the limits of text to speech: choosing from a small set of voices, adjusting scripts to coax out the right tone, and hoping the voice stays consistent through a longer exchange.

The launch of Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS marks a step forward. This is a new generation of text to speech that developers can direct with greater precision. Gemini 3.8 Flash TTS also brings an extended library of more than 2,000 voices, opening up far more choice across languages and regions.

At Agora, we’re excited about what this means for the people building real-time voice experiences. A more expressive voice gives an AI agent more ways to meet the moment, whether it’s helping a customer, guiding someone through an app, or bringing a story to life.

Control Speech Emotion, Pace, and Accent

Traditional text to speech starts with a script and produces audio. The words may be correct, but the delivery does not always match the intent. A welcome can sound flat. An apology can sound cheerful. A change in mood halfway through a sentence can be difficult to convey.

Gemini 3.8 Flash TTS gives developers more control over how speech is delivered. They can guide qualities such as emotion, pace, character, and accent, and shift the style as the moment changes. A character might begin a line quietly and end it with excitement. A voice agent might respond with warmth when a customer is frustrated, then sound more upbeat when the issue is resolved.

That control matters because conversation is dynamic. People adjust their tone as they listen and respond. Giving developers a way to shape those changes brings generated speech closer to the experience they are trying to create.

More Than 2,000 Voices Across Languages and Regions

Voice selection is another part of that creative process. Gemini 3.8 Flash TTS introduces an extended library of more than 2,000 voices, including options localized to specific languages and regions. The original named voices remain available as well.

For developers, that means more opportunity to find a voice that fits the experience and the people using it. A learning app, a game character, and a customer service agent may each call for a different sound. Regional voice options can also help a product feel more familiar to audiences in different places.

The size of the library is exciting, but the real value is choice. Teams can think beyond a handful of default voices and consider what their product should sound like for each audience and use case.

Voice Consistency and Multispeaker Speech

A convincing voice experience has to work beyond a single sentence. In a longer conversation, small changes in vocal quality or accent can become distracting. Exchanges between speakers also need to feel like a conversation, with natural pacing and responses.

Gemini 3.8 Flash TTS is designed to improve consistency through extended speech and to support more natural multi-speaker dialogue. It brings greater realism to elements such as breathing, pacing, and turn-taking. Those details can make a meaningful difference in an interactive story, a guided lesson, or a voice agent that spends several minutes helping someone complete a task.

For real-time applications, the goal is an experience people can stay engaged with. The voice should support the interaction, from the first greeting through the rest of the conversation.

What Developers Can Build with Gemini TTS

We see a broad range of possibilities for more expressive generated speech: voice agents that respond with appropriate tone, learning experiences that make dialogue engaging, characters that remain recognizable across a story, and applications that can speak to people in voices suited to their language and region.

Developers will decide what to build with Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. What makes this launch compelling is the additional room they now have to design the voice itself. They can think about delivery, identity, and the flow of a conversation alongside the words on the page.

At Agora, we believe the next wave of voice AI will be defined by experiences that feel natural to participate in. More expressive speech and a much wider range of voices give builders new tools to make those experiences possible.

← All articles

More in Media & Content

All →
Can Muse overcome Meta’s trust issues?Пресса
Meta

Can Muse overcome Meta’s trust issues?

Meta's Muse agent is attacking one of the economy's most profitable weak spotsПресса
Meta

Meta's Muse agent is attacking one of the economy's most profitable weak spots

Meta and YouTube say they will run ads for ‘Musk’ documentary after all
Пресса
Meta

Meta and YouTube say they will run ads for ‘Musk’ documentary after all

At Meta Connect, the company’s smart glasses were everywhereПресса
Meta

At Meta Connect, the company’s smart glasses were everywhere

Automattic has a new board after failed attempt to put CEO on leaveПресса
Automattic

Automattic has a new board after failed attempt to put CEO on leave

Meta opens early access program for new Muse featuresПресса
Meta

Meta opens early access program for new Muse features

More from Agora

Voice AI That Listens While It Talks: Inside GPT-Live-1
Agora

Voice AI That Listens While It Talks: Inside GPT-Live-1

Gemini 3.8 Live vs. Extended Thinking: Which Is Better for Voice AI?
Agora

Gemini 3.8 Live vs. Extended Thinking: Which Is Better for Voice AI?

How Turn-to-Turn Latency Shapes Voice AI?
Agora

How Turn-to-Turn Latency Shapes Voice AI?

Voice AI That Listens While It Talks: Inside GPT-Live-1
Agora

Voice AI That Listens While It Talks: Inside GPT-Live-1

Voice AI That Listens While It Talks: Inside GPT-Live-1
Agora

Voice AI That Listens While It Talks: Inside GPT-Live-1

Agora: turning real-time calls and streams into a competitive edge
Agora

Agora: turning real-time calls and streams into a competitive edge