AI Briefing
KOSign in

Samsung Labs Releases LittleBit-2 for Sub-1-Bit LLM Compression

·2026.10.08 22:29

Key point

The official implementation supports compression down to 0.1 bits per weight across models including Llama 3, Phi-4, and Qwen3.

Details

Samsung Labs has released the official code for LittleBit and its successor LittleBit-2, a method for compressing Large Language Models into the sub-1-bit regime. The project, accepted to NeurIPS 2025 and ICML 2026 respectively, utilizes latent factorization to binarize weight matrices while preserving model architecture at inference time.

Compression Methodology

The core technique factorizes dense weight matrices into low-rank latent factors, which are then binarized. Magnitude information is restored using lightweight learned scales, enabling extreme compression ratios of 1.0 to 0.1 bits per weight.

LittleBit-2 introduces an improvement called Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ). This step aligns SVD-derived latent factors with the binary hypercube during initialization to address geometry misalignment. This enhancement is opt-in via the --use_itq flag and adds no additional inference overhead.

Supported Models and Usage

The codebase supports a wide range of foundation models, including:

  • OPT
  • Llama and Llama 2/3
  • Phi-4
  • Qwen2.5 and QwQ
  • Gemma 2 and Gemma 3
  • Qwen3

The implementation supports Quantization-Aware Training (QAT) with SmoothSign and optional residual factorization. Users can train models using single GPU or multi-GPU setups with DeepSpeed. Evaluation scripts are provided for perplexity tasks (wikitext2, c4) and zero-shot tasks (boolq, piqa, hellaswag, etc.).

Technical Requirements

The project recommends Python 3.12 and CUDA 12.4.1. For reproducing paper results, it specifically requires transformers version 4.51.x, as newer releases may alter model internals. The repository is licensed under CC BY-NC 4.0.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.