Meta Unveils Multimodal Safety Model Llama Guard 4
Key point
Meta has released Llama Guard 4, a 12B-scale multimodal safety model that inspects both text and images.
Details
Llama Guard 4 is a 12B-scale dense multimodal model that analyzes both text and images to detect inappropriate content. It was optimized by removing the MoE (Mixture-of-Experts) layers from the Llama 4 Scout model and utilizing only the shared expert weights, and it can run on a single GPU with 24GB of VRAM.
This model can identify 14 types of hazards (violence, sexual content, privacy violations, election-related content, etc.) based on the MLCommons taxonomy, as well as code interpreter abuse. Compared to its predecessor, Llama Guard 3, performance has significantly improved in English and multimodal settings, and multilingual support is also available.
Additionally, the Llama Prompt Guard 2 series (86M, 22M parameters), specialized for detecting prompt injection and jailbreaks, was also released, enabling the construction of even more robust security pipelines.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.