1.58-bit LLM Training System for Ascend NPUs
·2026.05.25 00:24
Key point
Huawei unveiled the BitCPM-CANN model series, which supports 1.58-bit quantization-aware training on Ascend NPUs.
Details
BitCPM-CANN is the result of systematic research to perform 1.58-bit (ternary) quantization-aware training (QAT) on the Huawei Ascend NPU platform. The team ported the existing GPU-based pipeline to CANN, MindSpeed, and Megatron-LM, training 4 models at 0.5B, 1B, 3B, and 8B scales.
Key achievements include:
- Performance retention: The 1B, 3B, and 8B models retain 95.7%~97.2% of full-precision performance, with the 3B model showing comparable performance on the BBH benchmark.
- Maximized efficiency: Weight memory during inference can be reduced by up to 8x (about 6x end-to-end), while training throughput overhead is only 4.5%, enhancing practicality.
- Ecosystem contribution: Provides the first end-to-end 1.58-bit training system and infrastructure for the Ascend ecosystem, scalable up to 8B parameters.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.