SLAI T-Rex Optimizes DeepSeek-V4 Training with Ascend NPU
Key point
The Ascend NPU-based SLAI T-Rex framework achieved 34.22% MFU in post-training the trillion-parameter DeepSeek-V4 MoE model.
Details
SLAI T-Rex is an end-to-end optimization framework for full-parameter post-training of trillion-parameter-scale MoE models on Ascend NPU SuperPOD.
Key achievements:
- Achieved 34.22% MFU — a 2.93x improvement over the open-source baseline
- Hierarchical optimization covering the entire stack: model parallelism, compute-communication orchestration, and low-level kernel execution
An Operations Research (OR) specialized model was also released. Built on DeepSeek-V4-Flash with a CPT·SFT pipeline, it was trained on 10,000 high-quality SFT samples including solver-verified synthetic data.
- Achieved zero-shot Pass@1 of 71.81%
- Outperformed GPT-5.4-Mini by +3.98pp and the base DeepSeek-V4-Flash by +11.27pp
The code is available on GitHub, and the model is available on ModelScope.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.