Qwen3.8 2.4T Released
Key point
Qwen has released the open model Qwen3.8 with a total of 2.4T parameters and 95B active parameters.
Details
Qwen has released the weights and configuration files for Qwen3.8-2.4T-A95B, the top-tier generation of its open model product line, in Hugging Face Transformers format.
- Total parameters: 2.4T
- Active parameters: 95B
- Architecture: MoE activating 10 Routed Experts and 1 Shared Expert out of 512 Experts
- Context length: Base 262K tokens, expandable up to approximately 1.01M tokens
- Supported environments: Transformers, vLLM, SGLang, TokenSpeed, etc.
Qwen3.8 improves coding, professional tasks, research, and long-horizon agent tasks, allowing control over reasoning depth via reasoning_effort and maintaining reasoning context in conversation history via preserve_thinking.
In the evaluations presented in the model card, it recorded scores such as PaperBench 93.0, IFBench 82.8, and MRCR v2 256K 92.9, and also compared performance in coding agents, business automation, and tool use. A managed version offering vision input, non-thinking mode, a base 1M context, and official tools is provided separately as Qwen3.8-Max.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.