AI Briefing
KO

Qwen2.5-Coder: More Coding, More Training!

·2024.09.19 01:00

Key point

The next-generation open source model **Qwen2.5-Coder**, trained on 5.5 trillion tokens to dramatically strengthen coding and math capabilities, has been released.

1 / 2

Details

The next-generation open source coding model Qwen2.5-Coder has been released. As the successor to the existing CodeQwen1.5, it officially changes the model name to Qwen-Coder, presenting a vision of an even more agile coding partner. The model is available in 1.5B, 7B sizes, with a 32B version also to be released.

This update focused on scaling coding data while maintaining math and general capabilities. By training on 5.5 trillion tokens (including source code, text-code grounding data, and synthetic data), it has significantly boosted performance on coding-related tasks. At the same time, it reinforces math and general capabilities, laying a foundation suitable for practical applications such as Code Agent.

The key features of Qwen2.5-Coder are as follows:

  • Extensive support: Supports up to 128K context length and covers 92 programming languages.
  • Strong performance: The open source 7B version shows performance surpassing larger models such as DeepSeek-Coder-V2-Lite and CodeStral-22B.
  • Versatile capabilities: Beyond code generation, completion, and fixing, it maintains competitive performance on math and general benchmarks such as GSM8K and MMLU.

The instruction-tuned Qwen2.5-Coder-Instruct model demonstrates even greater versatility. It achieves outstanding results in multilingual evaluation via McEval, code reasoning based on CRUXEval, and mathematical reasoning capability, proving its 'scholarly' side as well.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.