The 1.58-bit shock
Key point
Ternary Bonsai has released 8B, 4B, and 1.7B models using 1.58-bit ternary weights.
Details
Ternary Bonsai has been released. Using {-1, 0, +1} ternary weights, it compresses down to 1.58-bit level, and it claims to reduce memory usage by about 9x compared to typical 16-bit models.
The model is provided in three sizes: 8B, 4B, and 1.7B. The announcement describes it as an extension of the recent 1-bit Bonsai, taking the approach of using slightly more capacity while minimizing performance degradation.
- Release form: Blog announcement and Hugging Face model collection
- Format support: Currently, MLX 2-bit packed format is the only packing format available
- Compatibility: 8B FP16 safetensors is also provided, allowing it to run with stock Hugging Face tools
- Performance claims: Introduced as outperforming most competing models of the same parameter class on standard benchmarks
This can be seen as an extension of the trend toward making ultra-low-bit LLMs practically usable in memory-constrained environments.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.