AI Briefing
KOSign in

Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

·2026.10.07 01:00

Key point

The 740M-parameter model unifies text, images, audio, and video into a shared embedding space under an Apache 2.0 license.

Details

Google has launched EmbeddingGemma 2, an open, lightweight multimodal embedding model designed for on-device inference. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, the model maps text, images, audio, and video into a unified embedding space. It features 740 million parameters and is optimized for privacy-first applications, enabling tasks like finding specific video clips from voice memos or searching audio recordings via text queries directly on local devices.

The model is modular, requiring as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support. It achieves leading scores among sub-1B multimodal embedders, including a 9.92-point improvement in code performance (rising from 68.76 to 78.68 on MTEB Code). To optimize storage, it uses Matryoshka Representation Learning (MRL), allowing vector truncation from 768 dimensions down to 512, 256, or 128, offering up to 6x storage reduction. On a Google Pixel 11 Pro, quantized weights require approximately 191MB of active RAM for text-only and 567MB for the full model.

EmbeddingGemma 2 features an 8K token context window (4x larger than its predecessor), supporting up to 5.5 minutes of audio, 29 images, or 58 video frames on-device. It shares its text tokenizer and audio encoder with Gemma 4, facilitating unified pipelines for on-device RAG. The model is available on Hugging Face and Kaggle, with support for tools like MediaPipe, LiteRT, WebGPU, Ollama, and LMStudio.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.