Kakao Enhances Kanana-o Performance: Strengthening Multimodal Instruction Following and Emotional Expressiveness
Key point
Emotional expressiveness improved through DPO training and actor recording data, surpassing GPT-4o on Korean image-based benchmarks
Details
Kakao's Kanana organization unveiled the evolution of its multimodal AI model Kanana-o. This update focuses on implementing AI that interacts naturally like humans in real-world service environments, rather than solely improving benchmark scores. While maintaining existing Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) performance, it significantly enhanced Multimodal Instruction Following capabilities and speech expressiveness.
Enhanced Multimodal Instruction Following Improved the ability to accurately interpret complex instructions by considering various modalities such as text, images, and audio. A high-quality dataset was directly constructed, including multi-step instructions for summarization, translation, and reasoning, as well as image-audio cross-reference data. As a result, it scored 91.22 on MIABench-Ko (Korean image-based), surpassing GPT-4o (89.51). However, it scored lower than GPT-4o and Gemini-2.5-Pro on some global benchmarks such as Speech AlpacaEval.
Strengthened Emotional Expressiveness and Natural Conversation Utilized DPO (Direct Preference Optimization) training and actual actor recording data to generate speech containing emotions and nuances beyond mere information delivery. The average score for emotional styles such as sadness and happiness, as well as tone expression, rose from 3.24 to 3.54. Additionally, multi-turn dialogue scenarios were expanded to handle unstructured requests such as podcast-style multi-speaker conversations, dialects, and context switching, enhancing the liveliness of conversations.
Future Roadmap Kakao plans to pursue Full-duplex voice conversation support and On-device model lightweighting. It will refine the structure to allow overlapping speech and immediate responses, while exploring hybrid architectures executable on smartphones and wearable devices. Furthermore, it plans to expand into a general-purpose audio model capable of generating music, environmental sounds, and Foley sounds, while strengthening safety through human preference alignment (RLHF).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.