Gemini 3.8 text-to-speech says hello

Gemini 3.8 text-to-speech says hello

Gemini 3.8 文本转语音功能正式发布

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 是我们迄今为止表现力最强的音频生成模型。用户可以在 Google AI Studio、Gemini API、Gemini Enterprise、Gemini Notebook 和 Google Vids 中生成自定义角色语音并指导场景对话。

Check out “Gemini 3.8 text-to-speech says hello” to see how our new models work. Create custom voices from scratch or replicate existing ones with simple natural language prompts. Direct your audio line-by-line to control pacing, emotion, and even realistic conversational sounds. Use these models for high-quality audiobooks, podcasts, or real-time voice agents at scale. We’ve included built-in safety tools like watermarking to keep your generated audio secure. 查看“Gemini 3.8 文本转语音功能正式发布”,了解我们的新模型如何运作。通过简单的自然语言提示,您可以从零开始创建自定义语音或复制现有语音。您可以逐行指导音频,以控制语速、情感,甚至是逼真的对话音效。这些模型适用于高质量有声读物、播客或大规模实时语音代理。我们还内置了水印等安全工具,以确保您生成的音频安全可靠。

Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids. 今天,我们为 Gemini 家族引入了两款全新的文本转语音模型,将语音生成从静态预设转变为动态创意工作室。这些模型使创作者、开发者和企业能够创造更丰富、更具表现力的音频体验,同时提升 Gemini Notebook 和 Google Vids 等产品的用户体验。

Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling. Gemini 3.8 Flash TTS: 专为深度创意指导和角色设计而打造。使用自然语言提示从零开始创建全新的语音,为游戏、沉浸式有声读物、播客和互动媒体中的角色注入生命。通过对表演提示、语速、方言转换和附和反馈的精细控制,逐行指导每一次表演。

Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance. Gemini 3.8 Flash-Lite TTS: 专为高容量、高性价比的规模化应用而打造。针对大规模配音、音频内容创作和富有表现力的语音代理进行了优化,并提供对语调、语速和情感细微差别的精细控制。

These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. 这些模型完善了我们快速增长的 Gemini 音频家族,此前该家族已包含 3.5 Live Translate、3.5 Transcribe、3.8 Live 和 3.8 Live Extended Thinking。

Create and customize your own voices

创建并自定义您自己的语音

Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences. 从 30 种原始语音扩展到无限的语音库。无论您需要全新的角色语音还是统一的品牌代言人,我们的 3.8 Flash TTS 模型都能为您提供一个完整的语音工作室。这使您能够为每个时刻创建并使用富有表现力、听起来自然的语音,同时赋能开发者和企业轻松构建自定义音频体验。

Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you’re bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. 生成式语音设计: 使用 Gemini 3.8 Flash TTS,通过自然语言提示,在 100 多种语言和方言中自定义角色、口音和语音特征,从零开始创建定制语音——无论您是要让一条戏剧性的喷火龙栩栩如生,还是塑造一位带有独特地区语调的魅力叙述者。

Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent. 语音复制: 只需 30 秒的语音样本(您自己的声音或您拥有使用权的语音),即可重新创建一致的语音配置文件。该功能由内置的同意验证、SynthID 水印和 C2PA 凭证支持,以保护开发者及其配音人才。

Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”). 语音混音: 即将推出。从我们的语音库中选择一个声音,并微调音色、音高、语速和口音。使用提示词来调整特征(例如“添加微妙的美国南部口音”或“使表达更柔和”)。

Direct the performance, line by line

逐行指导表演

Once you’ve selected your voices, both TTS models give you precise control over how each line is delivered. 一旦选定了语音,两款 TTS 模型都能让您精确控制每一行的表达方式。

Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene. 逐行指导表演: 编写您自己的舞台说明,或者让 Gemini 通过自然的脚本提示来引导表达——从冷静的客服人员到低声细语的悬疑场景,应有尽有。

Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks. 长文本生成: 在数小时的连续音频中保持高质量的语音、自然的语速和角色音色,且说话人特征偏移极小——非常适合播客和有声读物。

Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script —whether for a podcast or dramatic storytelling—while keeping both voices distinctly separated with natural conversational turn-taking. 原生双人场景编排: 通过单个脚本无缝指导多轮对话(无论是播客还是戏剧叙事),同时保持两种声音的清晰区分,并实现自然的对话轮流。

Scripted vocal bursts & backchanneling: Add realistic conversational texture using non-verbal cues (like , , ) and active-listening interjections (like |mhm| or |yeah|) for precise comedic timing and reaction beats. 脚本化语音爆发与附和反馈: 使用非语言提示(如 、、)和主动倾听的插话(如 |mhm| 或 |yeah|)添加逼真的对话质感,以实现精确的喜剧时机和反应节奏。