Mistral NeMo
Key point
Mistral collaborated with NVIDIA to release Mistral NeMo, a 12B model with enhanced Korean language performance.
Details
Mistral NeMo, developed by Mistral in collaboration with NVIDIA, is a 12B-scale model that offers a context window of up to 128k tokens. It delivers top-tier performance in its size class in terms of reasoning, common sense, and coding accuracy, and offers high compatibility that allows it to serve as a drop-in replacement for systems using Mistral 7B.
To help researchers and enterprises adopt it, both the pretrained base model and the instruction-tuned model have been released under the Apache 2.0 license. Additionally, thanks to quantization-aware training, FP8 inference is possible without performance degradation.
A new tokenizer called Tekken has been introduced to maximize compression efficiency for text and source code.
- Korean: 2x more efficient compression compared to before
- Arabic: 3x more efficient compression compared to before
- Others: About 30% improved efficiency for source code, Chinese, European languages, etc.
With strong multilingual support, it shows excellent performance across various languages including English, French, Chinese, Japanese, and Korean. Through advanced instruction fine-tuning, its ability to follow instructions, reason, engage in multi-turn conversations, and generate code has been significantly improved compared to Mistral 7B.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.