LongCat-2 Trained a 1.6 Trillion-Parameter Model on Chinese-Made Chips
Key point
LongCat-2 revealed that it trained a 1.6 trillion-parameter model using only 50,000+ Chinese-made AI ASICs.
Details
LongCat-2 (3.55TB, BF16) did not specify the chip vendor name in its official blog, but the page's meta description (in Chinese) states that "the entire training process was completed using domestic chips." This case demonstrates that training a frontier-class model is possible without Nvidia GPUs.
Key infrastructure figures:
- Pre-training performed with 50,000+ ASICs
- Trained on 35 trillion tokens or more, with no rollbacks or loss spikes
- Superpod = up to 48 chips connected via all-to-all high-bandwidth interconnect (similar to an NVLink domain)
- Connections between Superpods use a RoCE (RDMA over Converged Ethernet) fabric
- The two-tier design improved training throughput by approximately 30%
Addressing memory constraints: With less HBM capacity per chip compared to the H800 (80GB HBM), the team made active use of memory optimization techniques such as ZeRO-1 sharding, selective recomputation, and offloading of unused activations.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.