AI Briefing
KO

Google Unveils Lightweight Multilingual Embedding Model

·2025.09.04 09:00

Key point

Google has released EmbeddingGemma, a 308M-scale multilingual embedding model optimized for on-device environments.

Details

EmbeddingGemma, developed by Google DeepMind, is a high-performance small multilingual embedding model optimized for on-device use.

It features 308M parameters and a 2K context window, and supports over 100 languages. Among text-only multilingual embedding models under 500M in scale, it recorded top-level performance on the MTEB (Massive Text Embedding Benchmark).

It is designed based on the Gemma3 architecture, but uses bi-directional attention instead of the conventional causal attention to maximize its performance as an encoder model. This enables outstanding efficiency in retrieval tasks.

When quantized, it occupies less than 200MB of RAM, making it highly practical for resource-constrained environments such as mobile RAG pipelines and AI agents. It supports various frameworks including Sentence Transformers, LangChain, and LlamaIndex, and can also be fine-tuned on domain-specific data.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.