AI Briefing
KO

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

·2026.08.03 09:00

Key point

The study comparatively analyzed preference alignment methods and datasets for multimodal LLMs and proposed BDHS.

Details

Multimodal Large Language Models (MLLMs) suffer from hallucination, generating responses that do not match the image. Therefore, the core goal of preference alignment is to make the model's responses align more closely with image information.

The researchers independently analyzed the key components of MLLM alignment and categorized the algorithms into offline methods and online methods.

  • Offline methods: DPO (Direct Preference Optimization)
  • Online methods: Online DPO, PPO (Proximal Policy Optimization)

The analysis showed that in certain situations, combining offline and online methods could improve model performance. The study also explained that the construction method and detailed design of multimodal preference datasets have a significant impact on model performance.

The researchers also proposed BDHS (Bias-Driven Hallucination Sampling), a method for generating preference data without additional annotation or external models. BDHS creates data by leveraging the model's biased hallucination cases, and it showed competitive performance against existing multimodal alignment research across multiple benchmarks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.