Google Releases EmbeddingGemma 2
Key point
It integrates text, images, video, and audio into a 768-dimensional space, significantly improving code search performance.
Details
Google DeepMind has released EmbeddingGemma 2, which integrates text, code, images, video, and audio into a single 768-dimensional vector space. The architecture modularly combines a vision encoder (170M) and an audio encoder (300M) with the existing text-only model (270M), allowing users to load only specific modalities as needed to save resources.
Performance and Architecture
The context window has expanded from 2K to 8K tokens compared to the previous version, and it achieved a 9.92-point increase (68.76 → 78.68) on the code search benchmark (MTEB Code). Multilingual text performance remained similar to the previous version. While reducing the dimensionality to 128 maintains over 90% of text and code search performance, multimodal search performance drops sharply, so text-centric usage is recommended.
Deployment and On-Device Optimization
Licensed under Apache 2.0 for commercial use, it supports Hugging Face Transformers, PyTorch, TensorFlow, JAX, and more. On-device execution is optimized via LiteRT and MediaPipe; quantized, the text-only model uses 191MB of memory, while the full multimodal model uses 567MB. Text embedding latency is fast at 8.3ms on the Pixel 11 Pro TPU.
Precautions
Using float16 precision may cause quality degradation due to active value range overflow, so using bfloat16 or float32 is recommended. The pre-training data cutoff date is January 2025, and the Gemma Prohibited Use Policy must be followed.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.