AI Briefing
KO

Kakao Releases Open-Source Lightweight Multimodal AI 'Kanana-1.5-v-3b' Available for Commercial Use

·2025.07.24 00:00

Key point

Kakao has released the lightweight multimodal language model 'Kanana-1.5-v-3b' as open source for commercial use, achieving superior performance compared to competing models on Korean benchmarks.

1 / 9

Details

Kakao has released 'Kanana-1.5-v-3b', a lightweight multimodal language model licensed under the commercially usable Kanana License. This model features approximately 3.6 billion (3.6B) parameters and utilizes a ViT-based Vision Encoder alongside its proprietary C-Abstractor architecture. It demonstrated high performance compared to competing models on Korean-specific benchmarks (such as KoOCRBench), particularly scoring 85.93 on KoOCRBench, significantly surpassing the second-place Qwen2.5-VL (50.67). With an average score of 74.00 on English benchmarks, it has secured global competitiveness. The training process employed knowledge distillation (KD) techniques to transfer knowledge from a larger model (Kanana-1.5-v-9.8b), along with instruction following and Direct Preference Optimization (DPO) using high-quality Korean datasets (KoMIF, KoText-HQ). The model is available for download on Hugging Face, and Kakao is currently developing future models with enhanced reasoning capabilities and an integrated multimodal model, 'Kanana-o'.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.