HyperCLOVA X OMNI: Korea's National AI, the Journey Toward Omnimodal
Key point
Naver has unveiled an omnimodal AI connecting text, images, and audio with 32B Think and 8B Omni.
Details
The K-AI that Team Naver envisions is a sovereign AI that all citizens can easily use. Rather than simply a model that answers well, it aims to understand Korea's language, culture, and social context, solve real-life problems, and go beyond text to handle images and audio as well.
At the core of this trend is Omnimodal. This refers to a structure where a single model simultaneously understands and generates text, images, and audio, going beyond the level of existing VLMs that only explain images through text. Instead of attaching separate models for each modality, Team Naver seeks to integrate various senses within a single model to achieve more natural and stable interaction.
The focus of this release is two models.
- HyperCLOVA X SEED 32B Think: A structure where speech recognition and speech synthesis modules are attached before and after an existing text-image VLM
- HyperCLOVA X SEED 8B Omni: A true omni-model trained from the start with text, images, and audio as a single integrated model
32B Think serves as a bridge that maintains the reasoning capabilities of existing VLMs while also supporting voice conversation. When a user shows a photo and asks a question by voice, the model understands the image, answers in text, and then reads the answer back aloud. However, due to its structure, there may be output constraints and latency.
8B Omni, on the other hand, starts from a different point. Since it is trained by aligning the three modalities into a single semantic space rather than processing them separately, it aims to deliver consistent understanding and responses regardless of whether the input is text, images, or voice. Its more streamlined architecture also makes it easier to scale up to larger omni-models later.
The roles of the two models are also clearly distinguished in performance evaluation. 32B Think showed superior results compared to comparison models in text understanding/generation, visual understanding, and agentic task performance, leading especially in agentic task performance by a margin of over 15 percentage points. 8B Omni evenly handled various input-output combinations across 12 global multimodal benchmarks, demonstrating balanced performance rather than being strong in only specific combinations.
Real-world use cases were also presented. 32B Think achieved top-grade (Level 1) results in major subjects such as Korean, English, Math, and Korean History on the 2026 CSAT (Korea's college entrance exam), scoring full marks in English and Korean History. 8B Omni was applied to actual agents, demonstrating its omnimodal potential through examples such as consultation, multilingual/dialect voice conversion, and image style conversion.
Ultimately, this announcement is less a finished product and more a starting point for an omni-model that will expand into Team Naver's future agentic AI system. The direction toward an AI that naturally understands and generates not just text but also vision and hearing has become clear, and the remaining task now is to scale up this structure to build greater real-world problem-solving capability.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.