Intelligent transcription with Gemini 3.5 Transcribe

Intelligent transcription with Gemini 3.5 Transcribe

使用 Gemini 3.5 Transcribe 实现智能转录

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. 今天,我们正式推出 Gemini 3.5 Transcribe,这是我们迄今为止最精确的语音转文字模型,专为智能语音交互而设计。与传统语音识别模型在处理背景噪音、复杂术语和口语纠错方面表现不佳不同,Gemini 3.5 Transcribe 能将原始音频直接转换为准确、精炼且格式规范的文本。

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. 在 Gemini 应用和 Android 等产品中,我们已经看到用户通过 Android 上的 Rambler 和 macOS 版 Gemini 应用等新语音功能,从该转录模型中获益。现在,开发者可以通过 Google AI Studio 和 Gemini 企业代理平台(Gemini Enterprise Agent Platform)中的 Gemini API,利用 Gemini 3.5 Transcribe 构建类似的功能。

We’ve built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs: 我们构建 3.5 Transcribe 的初衷是使其能够无缝接入开发者的工作流,无论您是在构建语音代理、实时字幕工具,还是通话后分析管道。该模型通过两个独立的 API 提供:

  • Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.

  • 实时流式传输: 通过 Live API 使用 gemini-3.5-transcribe-live,为交互式语音应用提供亚秒级延迟的连续双向流式传输。

  • Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

  • 预录音频处理: 通过 Interactions API 使用 gemini-3.5-transcribe,对录音、会议、通话记录等进行转录,并提供说话人识别和词级时间戳。

Get more precise and intelligent transcription

获取更精确、更智能的转录

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice. Gemini 3.5 Transcribe 旨在捕捉您的自然说话风格,以更好地理解您的意图并识别自定义词汇,从而让您能够通过语音执行任务。

  • Smart transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler words (“ums” and ‘“ahs”), auto-formats your text.

  • 智能转录: 无缝处理自我修正(例如“我们周二见面——不,周三”),去除填充词(如“嗯”、“啊”),并自动格式化文本。

  • Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.

  • 函数调用: 该模型可以通过函数调用将复杂任务(如图像生成和文件分析)委托给其他 Gemini 模型。目前已在 Gemini macOS 应用中提供。

  • More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.

  • 更精确的转录: 根据 Artificial Analysis 的测量,其流式传输的平均词错误率(WER)为 4.0%,非流式传输用例的平均词错误率为 2.6%。它在嘈杂的现实环境中表现出色,能准确捕捉邮政编码和订单 ID 等字母数字实体。

  • Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.

  • 自定义词汇: 通过无缝适配您提供的自定义词汇,识别专业术语和独特的拼写。

  • Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.

  • 全球语言支持: 自动检测并转录超过 85 种语言,无缝处理地区口音和各种方言。

  • Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

  • 多说话人识别: 在预录音频中准确识别最多三名说话人,并提供时间戳(支持 3 人以上的功能目前处于实验阶段)。

Experience smart transcription and advanced dictation

体验智能转录与高级听写

In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease. 除了 Google AI Studio 和 Gemini 企业代理平台中的 Gemini API 外,3.5 Transcribe 的功能远超标准的语音转文字,使 Google 各项服务的使用体验更加自然直观。通过将上下文感知理解直接引入 Gboard、Antigravity、Gemini 应用和 Chrome 等日常界面,它能轻松捕捉细微差别、意图和行内编辑。

On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style. 在 Android 版 Gboard 上,通过全新的 Rambler 功能,3.5 Transcribe 能将口述想法转化为格式规范的文本,并过滤掉填充词。您还可以使用语音进行编辑、纠正拼写错误并更改写作风格。

On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents. 在 Google Antigravity 上,3.5 Transcribe 会在您授权的情况下,结合屏幕上下文和聊天记录,确保在文件名、代理思考过程和活动文档中实现精准的转录。

In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice. 在 macOS 版 Gemini 应用中,3.5 Transcribe 不仅能将您的自然语音转录为整洁的格式化文本,还能启用语音指令,与屏幕上下文无缝结合以驱动复杂的工作流。通过在后台调用其他 Gemini 模型来处理繁重任务,该模型让您只需动动嘴,就能轻松总结本地文件、在不同应用间重用文本,或在光标处直接生成图像。

Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice. 即将登陆 Chrome 的功能将允许您在任何网页输入框中通过语音输入——让您能够更自然、更轻松地通过语音听写回复、起草帖子或向 Chrome 中的 Gemini 发出指令。