Kakao Open-Sources 'Kanana Safeguard' Series of Korean-Specialized AI Guardrails
Key point
Kakao has open-sourced the 'Kanana Safeguard' series, consisting of three AI guardrail models specialized for Korean user environments, under the Apache 2.0 license.
Details
Kakao has released the 'Kanana Safeguard' series of AI guardrail models via Hugging Face under the Apache License 2.0, reflecting Korean word order, honorifics, and cultural context. The series is divided into three models based on the nature of risks and detection scope.
- Kanana Safeguard (8B): Detects seven types of harmful content, including hate, harassment, sexual content, crime, child sexual exploitation, suicide and self-harm, and misinformation, by considering both user inputs and AI responses.
- Kanana Safeguard-Siren (8B): Analyzes only user inputs to detect legal and policy risks such as adult verification, professional advice (medical/legal, etc.), personal information, and intellectual property rights.
- Kanana Safeguard-Prompt (2.1B): A lightweight model designed to detect malicious prompt attacks, such as Prompt Injection and Prompt Leaking.
These models are designed to return judgment results in a 'Single Token' format, increasing inference speed and minimizing resource usage. Additionally, they were trained on a high-quality Korean dataset (approximately 240,000 items in total) generated by professional labelers, and demonstrated improved precision and recall compared to overseas benchmark models through evaluation data across four difficulty levels (Pass Required, Easy, Hard, Challenge). Kakao emphasized that AI safety is a public value and released these models to contribute to the ecosystem.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.