AI for everyone in every language

AI for everyone in every language

让每个人都能用自己的语言使用 AI

We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed. 我们正在超越传统的文本翻译,致力于构建能够精准理解全球丰富、鲜活语言表达方式的模型。

Today, our technologies and products power everyday interactions in more than 300 languages, spoken by more than 7 billion people — representing 86% of the global population. Reaching this milestone is meaningful, but it also underscores work that is critical to our mission. For decades, technology has worked best for a handful of dominant languages, leaving thousands of living languages and dialects poorly represented or absent altogether from the digital world. 如今,我们的技术和产品支持着全球超过 300 种语言的日常交流,覆盖超过 70 亿人口,占全球总人口的 86%。达到这一里程碑意义重大,但也凸显了我们使命中至关重要的一项工作。几十年来,科技主要服务于少数几种主流语言,导致数以千计的鲜活语言和方言在数字世界中表现不足,甚至完全缺失。

When we launched Google Translate in 2006, our goal was simple: to break down the barriers between languages. Advances in AI have helped us bring that vision to more people, expanding Translate from a handful of languages to more than 250 today. But translating text isn’t enough. Technology needs to understand how people actually communicate in the real world. So we focus our research and development on building systems that honor cultural nuance and the richness of human language, enabling everyone to participate and be understood on their own terms. 2006 年我们推出 Google 翻译时,目标很简单:打破语言间的壁垒。AI 的进步帮助我们将这一愿景带给更多人,使翻译支持的语言从最初的几种扩展到如今的 250 多种。但仅仅翻译文本是不够的。科技需要理解人们在现实世界中是如何交流的。因此,我们将研发重点放在构建能够尊重文化细微差别和人类语言丰富性的系统上,让每个人都能以自己的方式参与交流并被理解。

Going from text to true understanding

从文本翻译迈向真正的理解

Historically, speech recognition systems followed a rigid, multi-step process: transcribing audio into text, processing that text, and then synthesizing it back into audio. While functional, this pipeline strips away the richest parts of human communication: tone, pacing, emotion, and context. People don’t speak in perfectly neat, grammatical sentences. We laugh, overlap, hesitate, and weave multiple languages together mid-sentence, like when we speak Spanglish or Hinglish. 从历史上看,语音识别系统遵循着僵化的多步骤流程:将音频转录为文本,处理文本,然后再将其合成为音频。虽然功能可行,但这种流程剥离了人类交流中最丰富的部分:语调、语速、情感和语境。人们说话时并不会使用完美、整洁的语法句子。我们会大笑、插话、犹豫,甚至在句子中间混合多种语言,比如我们说“西式英语”(Spanglish)或“印式英语”(Hinglish)时那样。

To capture this, we moved beyond text transcripts to native audio intelligence — training models like Gemini to process audio directly as is, while also grasping both sound and intent. These efforts include: 为了捕捉这些细节,我们不再局限于文本转录,而是转向原生音频智能——训练 Gemini 等模型直接处理原始音频,同时把握声音和意图。这些努力包括:

  • Fluid real-time dialogue tools: Today, Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and 2,000+ language pairs, naturally capturing code-switching and emotional cues along the way. 流畅的实时对话工具: 如今,Gemini 3.5 Live Translate 支持 70 种语言和 2000 多种语言对的实时口语翻译,能够自然地捕捉对话过程中的语码转换和情感暗示。
  • Gemini 3.5 Transcribe: This is our most precise speech-to-text model yet, turning raw audio into polished, formatted text, even in noisy environments or with complex jargon. It also powers features like Rambler on Android Gboard, which removes filler words, fixes grammar and punctuation, and lets you edit or rewrite with voice commands and switch seamlessly between languages. Gemini 3.5 Transcribe: 这是我们迄今为止最精确的语音转文本模型,即使在嘈杂环境或涉及复杂术语的情况下,也能将原始音频转化为润色后的格式化文本。它还支持 Android Gboard 上的 Rambler 等功能,该功能可以去除口头禅、修正语法和标点符号,并允许用户通过语音指令进行编辑或重写,还能在不同语言间无缝切换。

The 1,000 Languages Initiative

“千种语言”计划

AI is helping us break down language barriers at a scale that was previously unimaginable. But reaching more people in their preferred language means going beyond the languages where AI performs best today: Our goal is to support the world’s 1,000 most-spoken languages. To help make that possible, our Universal Speech Model — trained on 12 million hours of audio — used cross-lingual transfer learning, techniques that enable models to transfer what they learn from data-rich languages, to improve speech understanding in languages with far less training data. AI 正在帮助我们以过去无法想象的规模打破语言壁垒。但要用人们偏好的语言触达更多人,意味着必须超越目前 AI 表现最好的那些语言:我们的目标是支持全球使用人数最多的 1000 种语言。为了实现这一目标,我们的通用语音模型(Universal Speech Model)——基于 1200 万小时的音频训练而成——采用了跨语言迁移学习技术,使模型能够将从数据丰富的语言中学到的知识迁移过来,从而提升对训练数据较少的语言的语音理解能力。

Putting communities at the heart of language data

将社区置于语言数据的核心

Because the web disproportionately represents a few dominant languages, teaching AI to understand underrepresented languages required us to rethink how we gather data. The solution is local grassroots partnerships. This localized approach has driven three of our most ambitious open-data partnerships: 由于互联网上少数主流语言占据了不成比例的份额,要让 AI 理解那些代表性不足的语言,我们需要重新思考数据收集方式。解决方案是与当地基层建立合作伙伴关系。这种本土化方法促成了我们三个最雄心勃勃的开放数据合作项目:

  • WAXAL: Built with partners including Makerere University and Digital Umuganda, WAXAL is a large-scale, open speech dataset covering 27 Sub-Saharan African languages, capturing tonal variation and conversational rhythms often missing from traditional datasets. WAXAL: 与马凯雷雷大学(Makerere University)和 Digital Umuganda 等合作伙伴共同构建,WAXAL 是一个大规模开放语音数据集,涵盖 27 种撒哈拉以南非洲语言,捕捉到了传统数据集中经常缺失的声调变化和对话节奏。
  • Project Vaani: In partnership with the Indian Institute of Science (IISc) and Bhashini, Project Vaani is mapping India’s linguistic diversity through a region-anchored rather than language-anchored approach, collecting to date more than 30,000 hours of speech across 109 languages from more than 155,000 speakers. Vaani 项目: 与印度科学理工学院(IISc)和 Bhashini 合作,Vaani 项目通过“以地区为中心”而非“以语言为中心”的方法绘制印度的语言多样性地图,迄今已收集了来自 15.5 万多名发言者、涵盖 109 种语言的 3 万多小时语音数据。
  • Amplify Initiative: We teamed up with more than 1,600 local experts and 20 universities across four continents to contribute 15,000 multimodal data points capturing local nuance. Amplify 计划: 我们与四大洲的 1600 多名当地专家和 20 所大学合作,贡献了 1.5 万个捕捉当地细微差别的多模态数据点。

We’re also building on our work prioritizing open-source language innovation through our new tool Language Explorer. It’s an interactive tool that visualizes LinguaMeta, the world’s largest open-source language data repository. Recognized by Fast Company for design innovation, it continuously maps more than 7,000 spoken, written, and signed languages. 我们还通过新工具 Language Explorer 继续推进以开源语言创新为优先的工作。这是一个交互式工具,用于可视化展示全球最大的开源语言数据库 LinguaMeta。该工具因其设计创新获得了《快公司》(Fast Company)的认可,它持续绘制着全球 7000 多种口语、书面语和手语的地图。