AI Briefing
KO

Google unveils multimodal-capable Gemma 3

·2025.03.12 09:00

Key point

Google has released Gemma 3, a new open model supporting multimodality and a 128k context.

Details

Google has unveiled Gemma 3, a new series of open-weight models. The models are available in 1B, 4B, 12B, and 27B parameter sizes, including both Pre-trained and Instruction-tuned versions.

Key features are as follows:

  • Multimodal capability: The 4B, 12B, and 27B models are multimodal, capable of processing both text and images (1B is text-only).
  • Extended context: The 1B model supports a 32k context window, while the other models support up to 128k tokens.
  • Strong multilingual support: Models of 4B and above support over 140 languages in addition to English.

Technically, the models raise the base frequency of RoPE (Rotary Positional Embeddings) to 1M to efficiently handle long contexts, and process visual information via a SigLIP image encoder. Notably, Gemma-3-27B-IT recorded performance surpassing Gemini 1.5-Pro on benchmarks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.