Context length expanded to 1 million tokens
Key point
Qwen2.5-Turbo has expanded its context length to 1 million tokens and significantly improved inference speed and cost efficiency.
Details
Qwen2.5-Turbo has greatly expanded its context length from the previous 128k to 1 million (1M) tokens. This is a massive volume equivalent to about 1 million English words, 10 novels, or 30,000 lines of code.
In terms of performance, it achieved 100% accuracy on the Passkey Retrieval task, and scored 93.1 points on RULER, a long-text evaluation benchmark, surpassing GPT-4 (91.6). It also maintains competitiveness on par with GPT-4o-mini in short-sequence processing capability.
Key improvements are as follows:
- Improved inference speed: Through the Sparse attention mechanism, the time to first token (TTFT) for processing 1 million tokens was reduced from 4.9 minutes to 68 seconds, about a 4.3x reduction.
- Cost efficiency: It maintains a price of ¥0.3 per 1 million tokens, and can process 3.6x more tokens than GPT-4o-mini at the same cost.
It is currently available through Alibaba Cloud Model Studio, HuggingFace Demo, and ModelScope Demo.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.