【AI科技快讯 2026-10-07】Google DeepMind于10月6日正式开源EmbeddingGemma 2——一款740M参数的轻量级多模态嵌入模型,可将文本(含代码)、图像、视频帧、音频统一映射到768维共享向量空间,是目前端侧可用的最强多模态嵌入模型,号称在同尺寸下性能超越参数量翻倍的竞品。
核心要点:EmbeddingGemma 2基于Gemma 4解码器架构构建,主打端侧部署;原生支持四种模态及任意组合的混合输入(如「图片+文本描述」或「视频片段+音频」);在Google AI Edge推理框架中可在手机、笔记本上实时运行,单次嵌入延迟低于30ms;模型已在Hugging Face、Ollama、Vertex AI全面上架,权重完全开源(Apache 2.0);官方发布的MTEB多模态基准测试中,在检索、分类、聚类三项任务上击败体量两倍于自己的竞品。
应用场景:多模态RAG检索(用户上传截图+文字描述检索相关文档);跨模态搜索(用图片搜视频片段、用音频搜配图);本地照片/视频语义整理;Agent工具调用前的相似度匹配;移动端无网络情况下的语义理解。
行业影响:EmbeddingGemma 2延续了DeepMind「端侧Gemma生态」战略,与Gemma 4文本模型、Whisper类语音模型形成完整本地化套件,端侧AI能力正从单一模态走向全模态融合;开源策略直接对标OpenAI text-embedding-3-small和Cohere embed-v3,但凭借多模态和端侧优势在企业本地化部署场景具备明显吸引力。
参考来源:
1. Google DeepMind Blog: https://deepmind.google/blog/embeddinggemma-2-an-open-lightweight-multimodal-embedding-model/
2. Google Developers Blog: https://developers.googleblog.com/en/google-ai-edge-with-embeddinggemma-2/
3. Hugging Face模型页: https://huggingface.co/google/embeddinggemma-2
4. Google AI for Developers: https://ai.google.dev/gemma/docs/embeddinggemma
5. Ollama模型页: https://ollama.com/library/embeddinggemma-2
6. MarkTechPost: https://www.marktechpost.com/2026/10/06/google-deepmind-releases-embeddinggemma-2-a-740m-open-multimodal-embedding-model-built-on-gemma-4/
7. The Decoder: https://the-decoder.com/google-claims-embeddinggemma-2-outperforms-rival-embedding-models-twice-its-size/
8. Jetstream Blog: https://jetstream.blog/en/google-deepmind-embeddinggemma-2/







