Introducing Mistral 3
Key point
Mistral has announced the Mistral 3 series, its next-generation models released under the Apache 2.0 license.
Details
Mistral has announced its next-generation models, Mistral 3. This series consists of Ministral, small dense models at 14B, 8B, and 3B scale, and Mistral Large 3, a sparse mixture-of-experts (MoE) model with a total of 675B (41B active) parameters. All models are released under the Apache 2.0 license, supporting free use by the developer community.
Mistral Large 3 was trained from scratch using 3,000 NVIDIA H200 GPUs, and delivers performance on par with the best open-weight models on the market for general prompts. In particular, it boasts industry-leading performance in image understanding and in multilingual conversation for languages other than English and Chinese.
To improve the model's accessibility and efficiency, Mistral collaborated with NVIDIA, vLLM, and Red Hat on optimization. Through NVFP4 format checkpoints, the model can run efficiently on Blackwell NVL72 systems or 8×A100/H100 nodes, and also supports low-precision execution via TensorRT-LLM and SGLang.
- Mistral Large 3 ranked 2nd in the LMArena leaderboard's OSS non-reasoning category.
- A reasoning version of the model is also planned for future release.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.