OSCAR: 2-bit KV Cache Quantization Technique Released
·2026.06.10 04:00
Key point
OSCAR, a technique using spectral covariance-aware rotation to reduce quantization loss in KV Cache, has been released.
Details
OSCAR (RotationZoo), a 2-bit KV Cache quantization technique for improving LLM inference efficiency, has been announced. This technique uses the Offline Spectral Covariance-Aware Rotation method to minimize the information loss that occurs during quantization.
Key features and provided resources:
- Supported models: INT2 KV quantization weights provided for various models including Gemma-4-12B-it, Qwen3-32B, Qwen3-4B-Thinking
- Framework support: Includes code usable in llama.cpp and sglang environments
- Technical basis: The related research paper has been published on arXiv, supporting the technical validity
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.