Qwen3.8-27B 2-bit Quantization Released
·2026.08.21 02:19
Key point
Escha Labs released a 2-bit quantized model of Qwen3.8-27B, demonstrating performance equivalent to FP8.
Details
Escha Labs has released the Qwen3.8-27B-Escha-W2 model. With a size of 10.15GB, it achieved a single-stream inference speed of 82.6 tok/s on an RTX 5090.
It achieved approximately 100% of FP8 performance across an average of 8 benchmarks, notably scoring 86.81 on LiveCodeBench v6, higher than FP8's 85.16.
- Runtime: Uses a custom SGLang-based runtime
- Open Source: Both the model and runtime stack will be released under the Apache-2.0 license
- Compatibility: A GGUF version for llama.cpp support will also be provided soon
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.