AI Briefing
KO

Kakao Develops Kanana Nano Series Using Minitron-Based Compression Techniques

·2025.01.10 00:00

Key point

Kanana Nano-2.1B maintains 93% of Kanana Essence's Korean performance despite a 26% size reduction.

1 / 13

Details

The Kakao Kanana Alpha organization developed the cost-efficient Small Language Model (SLM) series Kanana Nano utilizing Pruning and Distillation techniques. Starting from the existing Kanana Essence model and adopting and advancing the Minitron methodology, they achieved similar or superior English performance compared to similarly sized Llama 3.2 and Gemma 2, while achieving overwhelmingly higher Korean performance.

Compression Strategy and Performance

Kanana Nano-2.1B, which reduced model size by 26%, maintained 93% of the original model's Korean performance. Unlike previous studies that limited compression to one or two steps, they achieved meaningful performance down to ultra-small models for mobile devices through 5-6 iterative rounds of Pruning and Distillation. In particular, the method of grouping and removing Query heads that share Key/Value in the Grouped-Query Attention (GQA) structure was effective, and a technique was applied to average and Tie input/output embeddings to address the issue where the proportion of embedding parameters increases as models become smaller.

Embedding Model Development

They also released a Korean embedding model that converted Decoder-only models into Encoders by applying the LLM2Vec methodology. After undergoing steps of bidirectional Attention activation, Masked Next Token Prediction, and Contrastive learning, the model showed superior performance compared to Qwen 2.5 and Llama 3.2 on Korean Retrieval tasks in the MTEB benchmark. Notably, embedding models based on Pruned Kanana Nano recorded generally higher scores.

Future Plans and Release

Currently, Kanana Nano-2.1B and its Instruct and Embedding models are scheduled to be released as open source in the future. Kakao plans to continue research on reinforcing data in insufficient domains such as code and mathematics, and on technologies to minimize information loss caused by Pruning.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.