AI Briefing
KO

Qwen3.8-27B-GSQ-RCO-GGUF: 27B LLM compressed to ultra-low capacity down to IQ2_XS in GGUF format

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

·2026.09.02 05:19

The Qwen3.8-27B model has been released in GGUF format using GSQ and RCO techniques. It offers various quantization versions from IQ2_S to IQ3_XXS, enabling the execution of large language models on limited hardware.

Mixed-precision quantization based on imatrix was applied to balance size reduction and performance retention. Files supporting MTP (Multi-Token Prediction) are also included to help improve inference speed.

It is compatible with major inference engines such as llama.cpp, vLLM, and Ollama. Released under the Apache 2.0 license, it has few restrictions for commercial use and custom pipeline construction.

HuggingFace
HuggingFace model

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

The original page has no description.

image-text-to-text

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.