EschaLabs Releases 2-bit Qwen3.6-35B-A3B Quantized Model
·2026.07.30 06:25
Key point
A 2-bit Qwen3.6-35B-A3B model and a dedicated runtime have been released, dramatically reducing size while maintaining FP8-level performance.
Details
EschaLabs has released a 2-bit Qwen3.6-35B-A3B quantized model developed to improve practicality in consumer GPU environments. This model combines model-aware fine-tuning with a recovery process to minimize the performance degradation that occurs with low-bit quantization.
Key Features and Performance:
- Efficiency: Achieves a disk size of 12.3GB and inference speed of 225 tokens per second on an RTX 4090.
- Hardware Requirements: Can run on a single 24GB consumer GPU (or 16GB with reduced context).
- Performance Retention: Records performance similar to or exceeding the FP8 baseline on key benchmarks such as MMLU-Pro (80.9) and MATH-500 (93.8).
Technical Details:
- Rather than standard GPTQ or AWQ methods, it uses a pipeline that requires a dedicated custom runtime.
- Supports the SGLang engine, providing an OpenAI-compatible API, concurrency, tool calling, and structured outputs.
- A ZML runtime that can run without a Python environment is also provided, with plans to support MLX, vLLM, and GGUF in the future.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.