Mistral 7B
Key point
Mistral AI released Mistral 7B, a 7.3B-parameter model that outperforms Llama 2 13B, under the Apache 2.0 license.
Details
Mistral AI has released Mistral 7B, its most powerful language model for its size. Despite being a 7.3B parameter model, it outperforms Llama 2 13B on all benchmarks and approaches the performance of Llama 1 34B on many metrics.
Key technical features are as follows:
- It applies Grouped-query attention (GQA) to increase inference speed.
- Sliding Window Attention (SWA) allows it to handle longer sequences at a lower cost.
Mistral 7B is provided under the Apache 2.0 license, allowing unrestricted use. It is also easy to fine-tune, and a chat model that outperforms the Llama 2 13B Chat model is provided as well.
It also excels in terms of efficiency. In reasoning and STEM fields (such as MMLU), Mistral 7B matches the performance of existing Llama 2 models that are more than 3 times its size, achieving both memory savings and improved throughput at the same time.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.