Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

谷歌全新的 AI 转录功能可自动剔除语音中的“嗯”、“啊”等赘词

We got a new Gemini Audio model while we’re still waiting for the overdue Gemini 3.5 Pro launch. 在等待已久的 Gemini 3.5 Pro 发布之际,我们迎来了一款全新的 Gemini 音频模型。

Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we’re still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June. 谷歌为 Gemini Audio 更新了全新的转录功能,能够自动识别专业术语及超过 85 种语言。Gemini 3.5 Transcribe 是 Gemini 家族的新成员,继 3.5 Live Translate 发布后推出。目前,我们仍在等待谷歌发布其承诺于六月推出的 Gemini 3.5 Pro 模型。

Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.” 谷歌表示,3.5 Transcribe “代表了我们此前转录模型 Chirp 3 的重大进步”,特别是在多语言表现和措辞错误率方面。据谷歌称,该转录模型允许用户“仅通过语音进行自然编辑”,并能自动格式化文本,剔除“嗯”、“啊”等填充词。

Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps. 用户可以向模型提供自定义词汇表,使 3.5 Transcribe 能够自动适应独特的拼写要求和专业术语,从而避免手动修改这些词汇。此外,它还能在预录音频中识别多达三位发言者,并提供精确到词的时间戳。

Alongside 3.5 Transcribe, Google also said that 3.5 Live and 3.5 Live Experimental updates will be coming to Gemini Audio today that build on the existing speech recognition tech powering Gemini’s voice chat mode. After we published this story, Google then reached out to say these additional models aren’t being launched yet, and didn’t provide a new launch date. According to the information Google previously provided, Gemini 3.5 Live is better at handling mid-sentence interruptions, language recognition, and live visual processing, while Gemini 3.5 Live Experimental goes further by narrating its progress step by step in real time while it tackles reasoning on more complex tasks. 除了 3.5 Transcribe,谷歌此前还表示 3.5 Live 和 3.5 Live Experimental 更新将于今日登陆 Gemini Audio,这些更新基于驱动 Gemini 语音聊天模式的现有语音识别技术。但在我们发布报道后,谷歌联系我们称这些额外模型尚未发布,且未提供新的发布日期。根据谷歌此前提供的信息,Gemini 3.5 Live 在处理句中中断、语言识别和实时视觉处理方面表现更佳,而 Gemini 3.5 Live Experimental 则更进一步,能够在处理复杂任务时实时逐步叙述其推理进度。

Gemini 3.5 Transcribe is rolling out starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. It’s also available for developers in public preview in the Gemini API via AI Studio and Antigravity. Google says that Chrome support is coming soon. Gemini 3.5 Transcribe 即日起面向所有 macOS Gemini 应用用户推出英语版本,并已在部分国家和语言的 Android 设备上通过 Rambler 听写功能上线。开发者也可通过 AI Studio 和 Antigravity 在 Gemini API 的公开预览版中使用该功能。谷歌表示,Chrome 浏览器的支持也即将到来。

Update, August 26th: Google also mentioned two Gemini 3.5 Live and 3.5 Live Experimental in information provided to The Verge prior to publication, but now says that only 3.5 Transcribe is being announced today. 更新(8 月 26 日):谷歌在发布前提供给《The Verge》的信息中曾提及 Gemini 3.5 Live 和 3.5 Live Experimental,但目前表示今日仅发布 3.5 Transcribe。