AI Briefing
KO

LG AI Research's First Public Vision Language Model, EXAONE 4.5

·2026.07.16 09:00

Key point

LG AI Research has released EXAONE 4.5, a multimodal model that understands text and images simultaneously, as an open-weight model.

1 / 2

Details

LG AI Research has unveiled EXAONE 4.5, which extends the language processing capabilities of the existing EXAONE 4.0 into the visual domain. This model adopts a Native Multimodal Pretraining approach that trains text and visual information together from the initial stage, maximizing multimodal understanding.

The key technical features are as follows:

  • Integrated Visual Encoder: A proprietary visual encoder has been integrated, applying GQA (Grouped Query Attention) to reduce computational load and optimize inference speed.
  • Multi-Token Prediction (MTP): Moving away from the conventional sequential token generation approach, the model predicts the next tokens simultaneously, improving inference speed by more than 1.5x compared to before.
  • Industry-Specialized Performance: Optimized for STEM (Science, Technology, Engineering, Mathematics) fields and Document Understanding, it also shows excellent performance in Korean and visual context understanding.

Based on large-scale datasets, EXAONE 4.5 achieves balanced performance in both language and vision domains, with strengths in complex document-based tasks required in real-world industrial settings.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.