Gemini-3.5-Transcribe
Gemini-3.5-Transcribe
Intelligent transcription with Gemini 3.5 Transcribe Gemini 3.5 Transcribe 智能转录
Aug 26, 2026 2026年8月26日
Our latest speech-to-text model designed for precise and intelligent real-time transcription. 这是我们最新的语音转文字模型,专为精准、智能的实时转录而设计。
Diego Melendo Casado, Senior Director, Engineering, Gemini Audio; Luke Leonhard, Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team. Diego Melendo Casado(Gemini Audio 工程高级总监)与 Luke Leonhard(Gemini Audio 幕僚长)代表 Gemini Audio 团队发布。
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. 今天,我们推出了 Gemini 3.5 Transcribe,这是我们迄今为止最精准的语音转文字模型,专为智能语音交互而设计。与传统语音识别模型在处理背景噪音、复杂术语和口语纠错方面表现不佳不同,Gemini 3.5 Transcribe 能将原始音频直接转换为准确、精炼且格式规范的文本。
Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. 在 Gemini 应用和 Android 等产品中,我们看到用户已经从该转录模型中受益,例如 Android 上的 Rambler 功能以及 macOS 版 Gemini 应用中的新语音功能。现在,开发者可以通过 Google AI Studio 中的 Gemini API 和 Gemini 企业代理平台(Gemini Enterprise Agent Platform),构建类似的强大功能。
We’ve built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs: 我们构建 3.5 Transcribe 的初衷是让其无缝融入开发者的工作流,无论您是在构建语音代理、实时字幕工具,还是通话后分析管道。该模型通过两个独立的 API 提供:
-
Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using
gemini-3.5-transcribe-live. -
实时流式传输: 通过 Live API 使用
gemini-3.5-transcribe-live,为交互式语音应用提供亚秒级延迟的连续双向流式传输。 -
Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using
gemini-3.5-transcribe. -
预录音频处理: 通过 Interactions API 使用
gemini-3.5-transcribe,对录音、会议、通话记录等进行转录,并提供说话人识别和词级时间戳。
Get more precise and intelligent transcription
获取更精准、更智能的转录体验
Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice. Gemini 3.5 Transcribe 旨在捕捉您的自然说话风格,以更好地理解您的意图并识别自定义词汇,从而让您能够通过语音执行任务。
-
Smart transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler words (“ums” and “ahs”), auto-formats your text.
-
智能转录: 无缝处理自我修正(例如“我们周二见面——不,周三”),去除填充词(如“嗯”、“啊”),并自动格式化文本。
-
Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.
-
函数调用: 该模型可以通过函数调用将复杂任务(如图像生成和文件分析)委托给其他 Gemini 模型。目前已在 macOS 版 Gemini 应用中可用。
-
More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
-
更精准的转录: 根据 Artificial Analysis 的测量,其流式传输的平均词错误率(WER)为 4.0%,非流式传输用例为 2.6%。它在嘈杂的现实环境中表现出色,能准确捕捉邮政编码和订单 ID 等字母数字实体。
-
Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
-
自定义词汇: 通过无缝适配您提供的自定义词汇,识别专业术语和独特的拼写。
-
Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.
-
全球语言支持: 自动检测并转录超过 85 种语言,无缝处理各种区域口音和方言。
-
Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).
-
多说话人识别: 准确识别预录音频中的说话人,并为最多三名说话人提供时间戳(支持 3 人以上为实验性功能)。
Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases. Gemini 3.5 Transcribe 的性能较我们之前的转录模型 Chirp 3 有了重大提升,提供了新功能、更低的词错误率以及显著改善的延迟。以 Artificial Analysis 的测量为例,最终转录时间缩短了 70%。在涵盖多种主流语言和地区的 FLEURS 基准测试中,该模型表现出精准的多语言能力,优于 Chirp 3,在流式模式下实现了 5.50% 的 WER,在非流式用例中实现了 5.04% 的 WER。
Experience smart transcription and advanced dictation
体验智能转录与高级听写
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease. 除了 Google AI Studio 和 Gemini 企业代理平台中的 Gemini API 外,3.5 Transcribe 超越了标准的语音转文字功能,使 Google 生态系统中的工作体验更加自然直观。通过将上下文感知理解直接引入 Gboard、Antigravity、Gemini 应用和 Chrome 等日常界面,它能轻松捕捉细微差别、意图和行内编辑。
-
On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
-
在 Android 版 Gboard 上,通过全新的 Rambler 功能,3.5 Transcribe 能将口述想法转化为格式规范的文本,并过滤掉填充词。您还可以使用语音进行编辑、纠正拼写错误并更改写作风格。
-
On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
-
在 Google Antigravity 上,3.5 Transcribe 在获得您许可的情况下,结合屏幕上下文和聊天记录,确保在文件名、代理思考过程和活动文档中实现精准的转录。
-
In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
-
在 Google AI Studio 中,您可以在 Build 模式下访问 3.5 Transcribe,通过语音即时构建代码应用。
-
In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
-
在 macOS 版 Gemini 应用中,3.5 Transcribe 不仅能将您的自然语音转录为整洁的格式化文本,还支持语音指令,这些指令可与屏幕上下文无缝结合,驱动复杂的工作流。通过在后台调用其他 Gemini 模型来处理繁重任务,该模型让您仅凭语音即可轻松总结本地文件、在不同应用间重用文本或在光标处直接生成图像。
-
Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.
-
即将登陆 Chrome,您将能够在任何网页输入框中通过语音输入——让您能够更自然、更轻松地通过语音口述回复、起草帖子或向 Chrome 中的 Gemini 发送提示。