Introducing Qwen2
Key point
Qwen has released Qwen2, its next-generation model with significantly enhanced performance and multilingual capabilities.
Details
Qwen2, an evolution from Qwen1.5, has been released. This series provides a total of 5 sizes of Pretrained and Instruction-tuned models: 0.5B, 1.5B, 7B, 57B-A14B, and 72B.
Key features are as follows:
- Enhanced Multilingual Capabilities: In addition to English and Chinese, the model was trained on 27 additional languages, including Korean and Japanese, to improve multilingual ability.
- Improved Performance: Performance in coding and mathematics has been significantly improved, achieving SOTA performance on numerous benchmarks.
- Long Context Support: The Qwen2-7B-Instruct and 72B-Instruct models support a context length of up to 128K tokens.
Technically, GQA (Group Query Attention) has been applied to all model sizes to increase inference speed and optimize memory usage. For smaller models, Tie Embedding is utilized for parameter efficiency.
In particular, Qwen2-72B demonstrated performance surpassing Llama-3-70B across various domains including natural language understanding, knowledge acquisition, coding, and mathematics, and achieved better performance than its predecessor, Qwen1.5-110B, with fewer parameters.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.