Status Update on Low-Bit (1-bit/Ternary) Models and Inference Optimization
Key point
This summary covers the latest updates on low-bit quantized models, including 1-bit, 2-bit, and Ternary models, and the status of optimized inference engines.
Details
This covers the latest release status of low-bit quantized models (1-bit, 1.58-bit, etc.) and optimization updates for inference engines that support them.
Key Model and Performance Updates
- Bonsai: Updates to 1-bit and 1.58-bit (Ternary) version models. A recent 27B model has been released, with added support for CUDA and Vulkan backends. Notably, throughput improved by 15–40% through the merging of CUDA optimization PRs.
- Maple-Preview (20B-A1B): Recorded very high inference speeds of 200+ t/s on Mac Mini M4 and 120+ t/s on iPhone.
- Mach-1-Additive-35B: Achieved performance of up to 120 t/s in consumer laptop environments.
- Other Models: Includes BitCPM-CANN (OpenBMB), Hy-MT1.5 (Tencent), Neutrino-8B (FermionResearch), and Pestle-27B-Ternary (Doses-AI) for medical research.
Inference Engines and Implementations
Each model utilizes optimized custom llama.cpp forks or the MLX library for efficient execution, providing accelerated performance across various hardware backends (CUDA, Vulkan, Metal, etc.).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.