Google unveils multimodal Gemma 4
·2026.04.03 03:28
Key point
Google DeepMind has released the new Gemma 4 series of open models supporting vision and audio.
1 / 2
Details
Google DeepMind has unveiled the new Gemma 4 model series under the Apache 2.0 license. The lineup consists of models in 2B, 4B, 31B sizes and a 26B-A4B Mixture-of-Experts (MoE) model.
Key Features and Technology:
- Multimodal capabilities: All models natively process video and images, and in particular the E2B and E4B models support native audio input for speech recognition.
- Maximizing parameter efficiency: The smaller models (E2B, E4B) apply Per-Layer Embeddings (PLE) technology. This assigns unique embeddings to each decoder layer, designed to deliver strong performance in on-device environments while keeping the actual parameter count low.
- High intelligence density: Google focused on maximizing performance relative to parameter count, delivering strong reasoning capabilities even in the smaller models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.