AI Briefing
KO

TRL Expands Support for VLM Alignment Techniques

·2025.08.07 09:00

Key point

HuggingFace's TRL library now supports the latest alignment techniques such as MPO and GRPO for VLMs.

Details

HuggingFace's TRL (Transformer Reinforcement Learning) library has updated its latest alignment techniques to maximize the performance of Vision Language Models (VLM).

This update focuses on going beyond the limitations of the existing DPO (Direct Preference Optimization) to extract richer signals and improve scalability.

  • MPO (Mixed Preference Optimization): Combines DPO's preference loss, BCO's quality loss, and SFT's generation loss to improve VLM's reasoning ability.
  • GRPO and GSPO: Group-based policy optimization methods that provide better scalability in the latest VLM environments.
  • RLOO and Online DPO: Support more efficient and scalable multimodal alignment.

In addition, native SFT (Supervised Fine-tuning) support for VLMs and vLLM integration have been added, helping developers train and deploy the latest multimodal models more easily and quickly.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.