Qwen3.8-27B NVFP4 Quantization Released
Key point
A checkpoint quantizing Qwen3.8-27B to NVFP4 using the QUASAR QAT algorithm has been released, maintaining BF16-level performance.
Details
A checkpoint fully quantizing the Qwen3.8-27B model to NVFP4 (W4A4) using a new QAT (Quantization-Aware Training) algorithm called QUASAR has been released. The model underwent a 2,446-step Quantization-Aware Distillation (QAD) process using the original BF16 model as the teacher, and inference is supported on NVIDIA Blackwell GPUs via vLLM.
While attention and GDN layers are typically maintained at higher precision to avoid quality degradation, this model with QUASAR maintains performance nearly identical to BF16. The model size is 19.7GB, approximately one-third of the original BF16 size (55.6GB), and it achieves higher accuracy with a smaller footprint than other NVFP4 quantized models.
Performance Comparison (GPQA-Diamond / AIME26):
- Qwen3.8-27B (BF16): 0.9141 / 1.0000
- QUASAR NVFP4: 0.9091 / 1.0000
- unsloth NVFP4: 0.8939 / 0.9778
- Inferact NVFP4: 0.8763 / 0.9667
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.