Gemini 3.8 text-to-speech

Gemini 3.8 Text-to-Speech

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 是我们迄今为止表现力最强的音频生成模型。您可以在 Google AI Studio、Gemini API、Gemini Enterprise、Gemini Notebook 和 Google Vids 中生成自定义角色语音并指导场景对话。

Check out “Gemini 3.8 text-to-speech says hello” to see how our new models work. Create custom voices from scratch or replicate existing ones with simple natural language prompts. Direct your audio line-by-line to control pacing, emotion, and even realistic conversational sounds. Use these models for high-quality audiobooks, podcasts, or real-time voice agents at scale. We’ve included built-in safety tools like watermarking to keep your generated audio secure. 查看“Gemini 3.8 文本转语音问候”以了解我们的新模型如何工作。通过简单的自然语言提示,您可以从零开始创建自定义语音或复制现有语音。逐行指导您的音频,以控制语速、情感,甚至是逼真的对话声音。将这些模型用于高质量有声读物、播客或大规模实时语音代理。我们内置了水印等安全工具,以确保您生成的音频安全可靠。

Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids. 今天,我们向 Gemini 家族引入了两款全新的文本转语音模型,将语音生成从静态预设转变为动态创意工作室。这些模型使创作者、开发者和企业能够创造更丰富、更具表现力的音频体验,同时提升 Gemini Notebook 和 Google Vids 等产品的用户体验。

Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling. Gemini 3.8 Flash TTS: 专为深度创意指导和角色设计而打造。使用自然语言提示从零开始创建全新的语音,为游戏、沉浸式有声读物、播客和互动媒体中的角色注入生命。通过对表演提示、语速、方言转换和反馈式对话(backchanneling)的精细控制,逐行指导每一次表演。

Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance. Gemini 3.8 Flash-Lite TTS: 专为高容量、高性价比的扩展需求而打造。针对大规模配音、音频内容创作和富有表现力的语音代理进行了优化,并能对语调、语速和表现细节进行精细控制。

These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. 这些模型完善了我们快速增长的 Gemini Audio 家族,继 3.5 Live Translate、3.5 Transcribe、3.8 Live 和 3.8 Live Extended Thinking 之后推出。

Create and customize your own voices

创建并自定义您自己的语音

Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences. 从 30 种原始语音扩展到无限的语音库。无论您需要全新的角色语音还是统一的品牌代言人,我们的 3.8 Flash TTS 模型都能为您提供一个完整的语音工作室。这使您能够为每一刻创建并使用富有表现力、听起来自然的语音,同时赋能开发者和企业轻松构建自定义音频体验。

Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you’re bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. 生成式语音设计: 使用 Gemini 3.8 Flash TTS,通过自然语言提示,在 100 多种语言和方言中自定义角色、口音和语音特征,从零开始创建定制语音——无论您是要让一条戏剧性的喷火龙栩栩如生,还是塑造一位带有独特地区语调的魅力叙述者。

Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent. 语音复制: 只需 30 秒的语音样本(您的声音或您拥有使用权的语音),即可重新创建一致的语音配置文件。该功能由内置的同意验证、SynthID 水印和 C2PA 凭证支持,以保护开发者及其配音人才。

Direct the performance, line by line 逐行指导表演

Once you’ve selected your voices, both TTS models give you precise control over how each line is delivered. 一旦选定了语音,两款 TTS 模型都能让您精确控制每一行的表达方式。

Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene. 逐行指导表演: 编写您自己的舞台指导,或者让 Gemini 通过自然的脚本提示来引导表达——从冷静的客服人员到低声细语的悬疑场景。

Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks. 长篇生成: 在数小时的连续音频中保持高质量的语音、自然的语速和角色音色,且几乎没有说话人漂移——非常适合播客和有声读物。

Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script — whether for a podcast or dramatic storytelling — while keeping both voices distinctly separated with natural conversational turn-taking. 原生双人场景编排: 从单个脚本无缝指导多轮对话——无论是播客还是戏剧叙事——同时通过自然的对话轮流,保持两种声音的清晰区分。

Scripted vocal bursts & backchanneling: Add realistic conversational texture using non-verbal cues (like , , ) and active-listening interjections (like |mhm| or |yeah|) for precise comedic timing and reaction beats. 脚本化语音爆发与反馈: 使用非语言提示(如 <笑声>、<叹气>、<喘息>)和主动倾听的插话(如 |嗯| 或 |是啊|)增加逼真的对话质感,以实现精确的喜剧时机和反应节奏。