AI Briefing
KO

Mixtral Expert Model

·2023.12.11 09:00

Key point

Mistral AI has released Mixtral 8x7B, an open-weight model that outperforms Llama 2 70B.

1 / 2

Details

Mistral AI has released Mixtral 8x7B, a high-performance Sparse Mixture-of-Experts (SMoE) model. Distributed under the Apache 2.0 license, this model outperforms Llama 2 70B on most benchmarks and shows performance equal to or higher than GPT-3.5.

Mixtral has 46.7B total parameters, but uses an efficient structure that employs only 12.9B parameters per token. By having the router network at each layer select 2 out of 8 expert groups for processing, it achieves high performance at the same speed and cost as a 12.9B model.

Key capabilities include the following:

  • Handles a 32k token context
  • Supports English, French, Italian, German, and Spanish
  • Excellent code generation capability
  • Can be instruction-following fine-tuned to achieve an MT-Bench score of 8.3

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.