AI Briefing
KO

Qwen3.8-27B 2-bit Quantization Released

·2026.08.21 02:19

Key point

Escha Labs released a 2-bit quantized model of Qwen3.8-27B, demonstrating performance equivalent to FP8.

Details

Escha Labs has released the Qwen3.8-27B-Escha-W2 model. With a size of 10.15GB, it achieved a single-stream inference speed of 82.6 tok/s on an RTX 5090.

It achieved approximately 100% of FP8 performance across an average of 8 benchmarks, notably scoring 86.81 on LiveCodeBench v6, higher than FP8's 85.16.

  • Runtime: Uses a custom SGLang-based runtime
  • Open Source: Both the model and runtime stack will be released under the Apache-2.0 license
  • Compatibility: A GGUF version for llama.cpp support will also be provided soon

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.