AI Briefing
KO

Qwen2.5-LLM: Expanding the Boundaries of LLMs

·2024.09.19 01:00

Key point

Alibaba unveiled the Qwen2.5 series, trained on 18 trillion tokens to substantially boost performance.

Details

Alibaba has unveiled the new Qwen2.5 series. This series includes 7 decoder-only dense models ranging from 0.5B to 72B parameters. In particular, it adds new 14B and 32B models suited for production environments, as well as a 3B model for mobile devices, building out a lineup tailored to user needs.

The key upgrades are as follows:

  • Expanded dataset scale: The pretraining dataset has been expanded from the previous 7 trillion tokens to up to 18 trillion tokens.
  • Enhanced coding and math capabilities: Qwen2.5-Coder and Qwen2-math technologies have been integrated, dramatically improving coding and mathematical reasoning performance.
  • Improved knowledge and preference alignment: MMLU benchmark scores have risen, and improvements in Arena-Hard and MT-Bench scores mean responses are generated that more closely match human preferences.

Additionally, instruction following capability and long-text generation (up to 8K tokens) have been strengthened, and most models support a context length of 128K. The ability to understand structured data such as tables and generate JSON output has also been improved, enhancing practical usability.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.