AI Briefing
KO

Pixtral 12B Released

·2024.09.17 09:00

Key point

Mistral AI released Pixtral 12B, which maintains text performance while offering powerful multimodal reasoning capabilities.

1 / 2

Details

Pixtral 12B is a natively multimodal model trained jointly on text and images, designed to replace Mistral Nemo 12B.

The model combines a 400M-parameter vision encoder trained from scratch with a 12B-parameter multimodal decoder. It supports variable image sizes and aspect ratios, and provides the flexibility to process multiple images simultaneously through a long 128K-token context window.

Key performance and features are as follows:

  • It achieves 52.5% on the MMMU reasoning benchmark, surpassing numerous larger-scale models.
  • It excels at chart and figure understanding, document question answering, multimodal reasoning, and Instruction Following.
  • It does not sacrifice performance on existing text-only benchmarks such as coding and math while boosting multimodal performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.