ISTA-DASLab Releases GSQ-RCO Quantized GGUFs for Qwen3.8-Flash-Next and Expert-Pruned Coder Build
Key point
The new quantized models achieve near-BF16 performance at 3.50 bpw, while the expert-pruned Coder build reduces the resident working set to 29.6 GB.
Details
ISTA-DASLab has released GSQ-RCO quantized GGUFs for Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 176.9B parameters. The release includes four standard quantized variants ranging from 2.40 to 3.50 bpw and a specialized Coder build that removes half of the model's routed experts.
Quantization Techniques
The release utilizes two primary optimization methods:
- GSQ (Gumbel-Softmax Quantization): A post-training scalar quantization method that learns grid assignments and group scales to close the gap between scalar and vector quantization at low bit-widths.
- RCO (Riemannian Constrained Optimization): Enforces exact memory budgets via gradient descent on task loss, used here to assign quantization types and select experts for retention.
Performance Metrics
The IQ3_S variant (3.50 bpw, 83.6 GB) matches or exceeds the BF16 base model on key benchmarks:
- AIME25: 100.00
- GPQA-Diamond: 92.93 (vs 91.92 for BF16)
- LiveCodeBench v6: 86.86 (vs 87.43 for BF16)
- Task Average: 93.26 (vs 93.12 for BF16)
The Q2_0 variant (2.40 bpw, 66.4 GB) achieves a zero-shot average of 78.00, surpassing the BF16 value of 76.94 at approximately one-fifth the size.
Expert-Pruned Coder Build
The Coder build removes 256 of 512 routed experts per layer, selected by RCO optimizing KL divergence. Retained weights remain at 3.5 bpw.
- Effective Bitwidth: The combined pruning and quantization result in an average of 1.89 bits per parameter of the original transformer.
- Memory Footprint: The total file size is 58.4 GB, but the resident working set is only 29.6 GB, allowing it to run on a single 32 GB accelerator.
- Capability Retention: SWE-bench Verified score is 75.60 (91.3% of BF16), and LiveCodeBench v6 is 86.28 (98.7% of BF16).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.