AI Briefing
KO

Qwen3.8-27B SOTA GGUF Released

·2026.08.29 06:46

Key point

IST-DASLab released a Qwen3.8-27B GGUF model with significantly improved performance over previous versions by applying GSQ and RCO technologies.

Details

IST-DASLab released Qwen3.8-27B GGUF models applying new quantization technologies: GSQ (Gumbel-Softmax Quantization) and RCO (Riemannian Constrained Optimization). These models can run without modification on major inference engines such as llama.cpp, Ollama, and LM Studio.

Core Technologies and Performance

  • GSQ: A learnable scalar quantization technique that reduces the performance gap between scalar and vector quantization in the 2-3 bit range.
  • RCO: Assigns quantization types based on per-tensor importance to minimize task loss within capacity constraints.

Benchmark Results

  • 3.00 bpw (10.1 GB): Recorded performance identical to the original BF16 model on AIME25 (100.00), and showed approximations within 1 point of the original model on GPQA-Diamond and LiveCodeBench v6.
  • 2.75 bpw (9.3 GB): Achieved an AIME25 score of 100.00, with a zero-shot average score (75.70) exceeding that of the BF16 original model (74.34).
  • Comparative Advantage: Showed performance improvements of +10.0 points on AIME25, +8.6 points on GPQA-Diamond, and +4.6 points on LiveCodeBench compared to Unsloth Dynamic (IQ2_S) at approximately 8.4 GB.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.