Qwen2.5-1M: Qwen Releases Models Supporting Up to 1 Million Token Context Length
Key point
Qwen has released the Qwen2.5-1M series, open-source models supporting up to 1 million tokens of context length.
Details
Qwen has released Qwen2.5-7B-Instruct-1M and Qwen2.5-14B-Instruct-1M, open-source models supporting up to 1 million tokens of context length. This release also includes an inference framework based on vLLM for efficient deployment.
The new inference framework integrates a sparse attention approach, improving the processing speed for 1 million token inputs by 3x to 7x compared to before. This allows developers to deploy models with large-scale context more efficiently.
In terms of performance, the Qwen2.5-1M series shows better performance than the existing 128K version, and in particular, the Qwen2.5-14B-Instruct-1M model proved itself to be a strong open-source alternative by surpassing GPT-4o-mini on several benchmarks. Additionally, while expanding long-context processing capability, performance on short-text tasks was maintained similarly to existing models, preserving core capabilities.
Key features:
- Open-source models: 7B and 14B instruct models provided
- High-speed inference: vLLM-integrated framework based on sparse attention
- Excellent performance: High information retrieval and comprehension capability in 1M token environments
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.