AI Briefing
KO

LG AI Research 183

·2026.07.16 09:00

Key point

LG AI Research has revealed the technical principles of EXAONE, a multimodal model that understands text and images bidirectionally.

1 / 2

Details

While existing vision AI has focused on image recognition, or has remained limited to one-directional learning like OpenAI's DALL-E, which converts text into images, EXAONE is a bidirectional multimodal model that also possesses the ability to generate text from images.

The core of this model lies in AugVAE (feature-Augmented Variational AutoEncoder) technology. This improves upon the existing VQ-VAE structure, and is characterized by placing a Visual Codebook component multiple times at each intermediate stage of image compression.

Through this structural differentiation, the following performance improvements were achieved:

  • Precisely distinguishing and memorizing image patterns regardless of size
  • Better organizing and arranging data features within the Latent Space
  • Enabling advanced mutual understanding between text and images

LG AI Research trained the model using a total of 250 million image-text pairs, and through this secured the ability to render even abstract concepts such as 'the feeling of spring' as images, going beyond simple object recognition.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.