EmbeddingGemma 2: An open, lightweight multimodal embedding model
EmbeddingGemma 2: An open, lightweight multimodal embedding model
EmbeddingGemma 2:一款开源、轻量级的多模态嵌入模型
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space. EmbeddingGemma 2 是目前设备端多模态嵌入领域性能最强的模型,能够将文本、图像、音频和视频的组合原生映射到一个统一的嵌入空间中。
We introduced EmbeddingGemma last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past our expectations. With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines. 我们去年推出了 EmbeddingGemma,旨在为高质量文本嵌入提供一种轻量级选择,帮助您的应用程序直接在消费级硬件上组织、搜索和关联信息。开发者社区的反应超出了我们的预期。该模型下载量已超过 2000 万次,开发者们利用它构建了更智能的设备端搜索工具和注重隐私的检索增强生成(RAG)流水线。
Today, we’re launching EmbeddingGemma 2, expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model. 今天,我们正式发布 EmbeddingGemma 2,它不仅限于文本,还将代码、图像、视频和音频统一在一个共享的嵌入空间中。EmbeddingGemma 2 基于 Gemma 4 架构构建,采用商业友好的 Apache 2.0 许可证发布,拥有 7.4 亿参数,非常适合设备端推理。它能够帮助用户从语音备忘录中找到特定的视频片段,或根据文本查询搜索数小时的录音,所有这些都由单一的原生多模态模型处理。
Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 is: EmbeddingGemma 2 采用与 Gemini 嵌入模型相同的技术,具备以下特点:
-
Best-in-class for its size: Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks. 同等规模下的最佳性能: 在 MTEB(大规模文本嵌入基准)代码测试和 MAEB(大规模音频嵌入基准)等基准测试中,在 10 亿参数以下的模型中处于领先地位,同时在文本、视觉和音频任务中表现优于许多规模更大的模型。
-
Modular by design: Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal support. 模块化设计: 仅处理文本任务时最少只需 2.7 亿参数,并可通过可选的视觉(1.7 亿参数)和音频(3 亿参数)编码器实现完整的多模态支持。
-
Storage-efficient: Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions. This provides up to 6x storage reduction for local vector databases and memory usage. 存储高效: 通过使用 Matryoshka 表示学习(MRL),开发者可以将输出向量从 768 维动态截断至 512、256 或 128 维。这可为本地向量数据库和内存占用带来最高 6 倍的存储空间缩减。
-
Optimized for on-device performance: Runs efficiently within tight resource constraints. With quantization, on a Google Pixel 11 Pro, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model. 针对设备端性能优化: 在严格的资源限制下高效运行。通过量化技术,在 Google Pixel 11 Pro 上,EmbeddingGemma 2 的纯文本权重仅需约 191MB 运行内存,完整多模态模型仅需约 567MB。
-
Extended context ready: Features an 8K token context window (4x larger than EmbeddingGemma 1), allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware. 支持扩展上下文: 具备 8K token 的上下文窗口(比 EmbeddingGemma 1 大 4 倍),允许直接在本地硬件上处理长达 5.5 分钟的音频、29 张图像、58 帧视频或它们的交错组合。
Achieving top-tier quality for code, vision, and audio
在代码、视觉和音频方面实现顶级质量
EmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size. Find full evaluation metrics and model information in the EmbeddingGemma 2 model card. EmbeddingGemma 2 在保持 EmbeddingGemma 强大的多语言文本性能的同时,在代码性能上实现了 9.92 分的显著提升(在 MTEB 代码基准中从 68.76 提升至 78.68),非常适合本地代码库索引、语义代码搜索和编码智能体检索。在图像、视频、文档和音频方面,它为 10 亿参数以下模型树立了“单位参数质量”的新标准,甚至超越了一些规模为其两倍以上的专业模型。完整的评估指标和模型信息请参阅 EmbeddingGemma 2 模型卡。
Enabling semantic search, routing, and retrieval, fully on-device
实现完全在设备端的语义搜索、路由和检索
EmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline. EmbeddingGemma 2 将强大的功能直接带到了边缘硬件上。在本地生成嵌入有助于确保数据隐私,降低流水线延迟,并使开发者能够构建完全离线工作的跨模态搜索和检索系统。
When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint. 当与 Gemma 4 等生成式模型配合使用时,EmbeddingGemma 2 可实现能够理解复杂多模态数据的设备端 RAG 流水线。由于 EmbeddingGemma 2 基于 Gemma 4 构建并共享其文本分词器和音频编码器,开发者可以在统一的流水线中同时运行这两个模型,且总内存占用更低。
- Use text or an image to find the top matches in your media library based on semantic similarity. Try it in Google AI Edge Gallery’s Instant Media Search. 使用文本或图像,根据语义相似度在媒体库中查找最佳匹配项。请在 Google AI Edge Gallery 的“即时媒体搜索”中尝试。
- Locate specific moments in video using text or audio queries. Try it in Google AI Edge Gallery’s Video Moments Finder. 使用文本或音频查询定位视频中的特定时刻。请在 Google AI Edge Gallery 的“视频时刻查找器”中尝试。
- Pair EmbeddingGemma 2 for local file retrieval with Gemma 4 for contextual reasoning. Try it in the Google AI Edge Foresight app. 将 EmbeddingGemma 2 用于本地文件检索,并结合 Gemma 4 进行上下文推理。请在 Google AI Edge Foresight 应用中尝试。
- Create real-time decision engines leveraging multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API. 通过 MediaPipe Decision Task API,利用多模态上下文创建用于分类、路由和预测功能的实时决策引擎。
To learn how to build on-device search and RAG systems with LiteRT, read the Google AI Edge blog post. 要了解如何使用 LiteRT 构建设备端搜索和 RAG 系统,请阅读 Google AI Edge 博客文章。
Getting started with EmbeddingGemma 2
EmbeddingGemma 2 入门
We worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build: 我们与以下合作伙伴密切合作,确保 EmbeddingGemma 2 能在您的开发环境中即刻使用:
- Download the models: Find the model weights on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. Visit LiteRT Community on Hugging Face for models optimized for on-device. 下载模型: 在 Hugging Face 和 Kaggle 上查找模型权重,Gemini Enterprise Agent Platform Model Garden 也即将上线。访问 Hugging Face 上的 LiteRT 社区获取针对设备端优化的模型。
- On-device deployment: Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks or LiteRT for custom model integration. Build for the browser with transformers.js or WebGPU. 设备端部署: 使用 Google AI Edge MediaPipe 开发跨平台应用,实现开箱即用的嵌入、检索和决策任务;或使用 LiteRT 进行自定义模型集成。通过 transformers.js 或 WebGPU 构建浏览器端应用。
- Use your favorite development tools: Serve the model efficiently using transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio. Store your embedding vectors with Qdrant. 使用您喜爱的开发工具: 使用 transformers、sentence-transformers、MLX、vLLM、llama.cpp、SGLang、Ollama 和 LMStudio 高效部署模型。使用 Qdrant 存储您的嵌入向量。
- Fine-tuning: Follow guidance by Unsloth for how to fine-tune EmbeddingGemma 2 for your use cases. 微调: 遵循 Unsloth 的指南,了解如何针对您的用例微调 EmbeddingGemma 2。
Explore our developer guide, documentation, and guides for inference and fine-tuning. 探索我们的开发者指南、文档以及推理和微调指南。