40% Savings with 2-Bit Quantization
Key point
Using Hadamard rotation and a Lloyd-Max codebook, the PLI of Gemma 4 E2B was compressed to 2 bits.
Details
The per-layer input (PLI) embedding of Gemma 4 E2B accounts for 60.6% of total weights, taking up a large amount of storage space.
TurboQuant-H simplifies TurboQuant's idea to fit offline weight quantization, using Hadamard rotation instead of random orthogonal rotation, and performing 2-bit quantization with a per-group Lloyd-Max codebook. The QJL correction was removed.
The key settings are as follows.
- vocab size: 262,144
- embedding dim: 8,190
- group size: 128
- codebook: 4 centroids per group
- codebook overhead: 0.125 bits/element
- effective bit rate: 2.125 bits/element
As a result, the PLI weights were reduced from 2,496 MB to 624 MB, and the total model storage size decreased from 4,790 MB to 2,918 MB. Perplexity increased by only 0.06, from 1.8547 to 1.9111, with no measured slowdown.
The evaluation conditions were 128 self-generated WildChat-1M completions and 24,438 scored tokens, compared against the HF BF16 baseline of 1.2892 PPL. Additionally, in group size experiments, 128 was presented as the balance point between overhead and quality.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.