Qwen2.5: A Feast of Foundation Models
Key point
Alibaba unveiled the Qwen2.5 model series, trained on 18 trillion tokens, delivering powerful performance.
Details
Alibaba has released the Qwen2.5 model series, built on the latest large-scale dataset. This release includes the general-purpose language model Qwen2.5, along with the coding-specialized model Qwen2.5-Coder and the math-specialized model Qwen2.5-Math.
The Qwen2.5 LLM was pretrained on up to 18 trillion (18T) tokens, achieving a significant boost in knowledge, coding, and math capabilities compared to the previous version, Qwen2. Key features include:
- Improved performance: Achieves MMLU 85+, HumanEval 85+, MATH 80+
- Scalability: Supports up to 128K context and can generate up to 8K tokens
- Multilingual support: Supports 29+ languages, including Korean
- Structured output: Enhanced ability to understand and generate structured data such as JSON
The specialized model Qwen2.5-Coder was trained on 5.5 trillion tokens of code data, delivering strong performance despite being a small model, while Qwen2.5-Math integrates multiple reasoning approaches such as Chain-of-Thought (CoT) and Program-of-Thought (PoT).
The largest model, Qwen2.5-72B, shows performance on par with Llama-3.1-70B and Mistral-Large-V2, and its base model achieves a level of performance competitive with massive models like Llama-3-405B.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.