AI Briefing
KO

Kakao Unveils Multimodal LLM 'Kanana-o' Integrating Voice and Vision via Model Merging

·2025.05.01 00:00

Key point

Kakao has released 'Kanana-o', a multimodal LLM that integrates its voice model (Kanana-a) and vision model (Kanana-v) using model merging techniques, demonstrating superior performance compared to global models in Korean speech recognition (CER 8.5).

1 / 9

Details

The Kakao Kanana organization has released 'Kanana-o', an integrated multimodal language model capable of understanding and generating text, images, and audio. This model was developed by leveraging the fact that the previously developed voice-specialized model (Kanana-a) and vision-specialized model (Kanana-v) share the same LLM backbone, combining the two models using Model Merging techniques and undergoing short additional training. This approach achieved performance equivalent to training from scratch while securing development time and flexibility. In performance evaluations, it recorded a score of 8.5 in Korean speech recognition (CER), overwhelming global models such as GPT-4o (17.5), and also showed competitive results in emotion recognition and image-audio integrated understanding benchmarks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.