AI Briefing
KO

Qwen1.5-110B: The First 100B+ Model in the Qwen1.5 Series

·2024.04.25 14:33

Key point

The Qwen team has released Qwen1.5-110B, a 100B-scale model with performance on par with Llama-3-70B.

Details

The Qwen team has released Qwen1.5-110B, the first model in the Qwen1.5 series with more than 100 billion parameters. This model shows performance on par with Meta-Llama3-70B in base model evaluations, and also achieved excellent results in chat evaluations such as MT-Bench and AlpacaEval 2.0.

Qwen1.5-110B is built on the same Transformer decoder architecture as the existing Qwen1.5 models. Its key features are as follows.

  • Application of Grouped Query Attention (GQA) for efficient model serving
  • Support for context lengths of up to 32K tokens
  • Extensive multilingual support, including Korean, English, Chinese, and Japanese

According to benchmark results, the 110B model shows performance competitive with Llama-3-70B and Mixtral-8x22B on key metrics such as MMLU, GSM8K, and HumanEval. In particular, compared to the existing 72B model, the performance improvement from scaling up model size is clearly evident.

Chat model performance also showed meaningful improvement over the existing 72B model. This suggests that, even without significantly changing the post-training approach, a stronger and larger base model can lead to a better chat model.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.