AI Briefing
KO

The New Hwadamsup: Image Understanding and Generation via LMM

·2026.07.16 09:00

Key point

LG AI Research has proposed the EXAONE LMM architecture, which resolves data licensing issues in existing LMMs while improving efficiency.

1 / 2

Details

Recently, AI technology has been evolving into a partner that assists artistic creative processes, and The New Hwadamsup Project is a representative example of media art realized through human-AI collaboration. This project utilized LG AI Research's EXAONE image understanding and generation technology.

Image Understanding is the task of answering questions about images using a VLM (Vision-Language Model) or LMM (Large Multimodal Model). A representative model, LLaVA, has a structure that converts images into a form an LLM can understand through a vision encoder and a projection layer.

However, existing LMMs rely on massive amounts of GPT-generated Instruction Tuning data to improve performance, which carries licensing issues and high cost problems. The EXAONE LMM was proposed to solve this.

The key features of EXAONE LMM are as follows:

  • Structural Differentiation: After extracting image information through Captioning and Detection modules, a SLM (Small Language Model) organizes this into a single sentence and passes it to the LLM.
  • Efficient Training: Instead of directly training the LLM, only the SLM, which handles organizing the information, is trained, enabling training with far less data and at much faster speed.
  • Improved Precision: By leveraging high-level vision features, it shows superior performance on tasks such as object Counting, an area where existing models were weak.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.