Pick
NVIDIA Releases Qwen3.6 NVFP4
·2026.05.31 02:49
Key point
NVIDIA has released a quantized version of Alibaba's Qwen3.6-35B model using the NVFP4 data type.
Details
NVIDIA has released a quantized version of Alibaba's Qwen3.6-35B-A3B model using the NVFP4 data type.
This model applies Post Training Quantization using NVIDIA Model Optimizer, and is optimized for inference via vLLM.
Key Features:
- Memory Reduction: By reducing bits per parameter from 16-bit to 4-bit, disk and GPU memory requirements were reduced by approximately 3.06x.
- Quantization Scope: Quantization was applied to the weights and activation functions of linear operators in the transformer blocks within the MoE (Mixture of Experts) architecture.
- Performance Retention: The model maintained high accuracy with very minimal performance degradation compared to BF16 on major benchmarks including MMLU Pro (85.0) and GPQA Diamond (94.7).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.