AI Briefing
KO

Kakao Unveils 'Kanana Safeguard' Series of Korean-Specialized AI Guardrails

·2025.09.19 00:00

Key point

Kakao has open-sourced three models under Apache 2.0 that detect harmful content, legal risks, and prompt attacks.

1 / 3

Details

Kakao has unveiled the Kanana Safeguard series, developed to ensure the safety of generative AI. This project consists of Korean-specialized guardrail models designed to address issues where AI generates harmful responses or becomes vulnerable to malicious prompt attacks.

Three Core Models and Their Roles

The Kanana Safeguard series comprises three models based on the nature of the risk and the scope of detection.

  • Kanana Safeguard (8B): Analyzes user inputs and AI responses to detect seven types of harmful content risks, including hate speech, sexual content, and criminal activity.
  • Kanana Safeguard-Siren (8B): Proactively blocks legal and policy risks at the user input stage, such as violations of the Youth Protection Act and the need for medical or legal advice.
  • Kanana Safeguard-Prompt: A lightweight model focused on input text classification to filter malicious prompt attacks.

Korean Specialization and Lightweight Strategy

To overcome the difficulty existing English-based guardrail models face in capturing Korean word order, honorifics, and metaphors, Kakao built a high-quality Korean dataset using professional labelers and LLM-synthesized data. Notably, the Kanana Safeguard-Prompt model was developed as a small model (2.1B class) considering the trade-off between cost and performance, adopting a strategy of improving performance through continuous monitoring after deployment. Additionally, all models are designed to process output values as a single token to increase inference speed and reduce resource consumption.

Open Source Release

In May 2025, Kakao released the Kanana Safeguard series under the Apache License 2.0 via Hugging Face. This initiative aims to treat AI safety as a public value and contribute to building a trustworthy AI ecosystem for developers and service providers.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.