AI Briefing
KO

MixAtlas: Uncertainty-Aware Data Mixture Optimization for Multimodal LLM Mid-Training

·2026.04.16 09:00

Key point

It optimized multimodal data mixtures at 1/100 the cost using a small proxy model and Gaussian process.

Details

MixAtlas is a framework that systematically optimizes data mixtures in multimodal LLM mid-training using an uncertainty-aware approach, rather than simply tuning them by intuition. While existing methods adjusted mixture ratios by looking at only one perspective, such as data format or task type, this method decomposes data along two axes—image concepts and task supervision—to enable more interpretable mixture control.

It leverages a small proxy model and a Gaussian-process surrogate to explore the mixture space at roughly 1/100 of the total training cost. The mixture ratios found this way transfer to actual large-scale training, preserving both efficiency and accuracy gains together.

The results are clear.

  • It showed up to 3x faster convergence compared to existing methods.
  • It recorded a consistent 2-5% performance improvement across multiple benchmarks.
  • It was particularly strong on text-heavy benchmarks: ChartQA +10%, TextVQA +13%.

The authors emphasize that this approach makes multimodal data mixture optimization both practical and interpretable. It was also accepted at the NADPFM at ICLR 2026 workshop, presented as a concrete recipe for training next-generation MLLMs.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.