AI Briefing
KO

Introducing the Qwen Series

·2024.01.23 23:13

Key point

The **Qwen** series, aiming at AGI, has been released, including open-source **LLMs** and **LMMs** of various sizes.

1 / 2

Details

Qwen is a project that goes beyond a simple language model to aim for AGI (Artificial General Intelligence), encompassing both LLMs (Large Language Models) and LMMs (Large Multimodal Models). Based on the base model Qwen, it consists of the chat model Qwen-Chat, the coding-specialized Code-Qwen, the math-specialized Math-Qwen, and the vision and audio models Qwen-VL and Qwen-Audio.

Currently, models of various sizes such as 1.8B, 7B, 14B, 72B have been released, of which 4 are provided as open source. All models were pretrained on 2-3 Trillion tokens, and are multilingual models that support various languages while showing strong performance in English and Chinese.

Key features are as follows:

  • Most open-source models support an extended context length of 32K.
  • An efficient Tokenizer with a high compression ratio is used.
  • Qwen-72B shows performance competitive with Llama 2, GPT-3.5, and GPT-4.

After pretraining, an Alignment process through SFT (Supervised Fine-Tuning) and RLHF (Reinforcement Learning from Human Feedback) is carried out to build high-quality chat models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.