Qwen2.5-Max: Exploring the Intelligence of a Large-Scale MoE Model
Key point
Alibaba unveiled **Qwen2.5-Max**, a large-scale MoE model trained on over 20 trillion tokens.
Details
Qwen2.5-Max is a large-scale MoE(Mixture-of-Experts) model developed through pretraining on over 20 trillion tokens, followed by sophisticated SFT(Supervised Fine-Tuning) and RLHF(Reinforcement Learning from Human Feedback) processes.
In performance evaluations, Qwen2.5-Max recorded performance surpassing DeepSeek V3 on major benchmarks such as Arena-Hard, LiveBench, LiveCodeBench, and GPQA-Diamond. It also showed competitive results on MMLU-Pro.
In Base model comparisons, it also demonstrated superior performance exceeding major models such as DeepSeek V3, Llama-3.1-405B, and Qwen2.5-72B.
The model can currently be experienced directly through Qwen Chat, and is provided as an API compatible with OpenAI-API through Alibaba Cloud.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.