AI Briefing
KO

Zai Optimizes GLM-5.1 Inference with ZCube Adoption

·2026.05.28 22:09

Key point

Zai introduced the ZCube network architecture, cutting GLM-5.1 inference costs by 33% and boosting throughput by 15%.

Details

Zai upgraded its existing ROFT network structure to ZCube for GLM-5.1 coding inference running on a 1,000-GPU cluster. ZCube is an architecture co-developed with Tsinghua University and HarnetsAI.

This upgrade achieved the following results:

  • 33% reduction in switch and optical module costs
  • 15% increase in GPU inference throughput
  • 40.6% reduction in P99 tail latency for the first token

While the existing ROFT topology was well-suited for training workloads, it suffered from asymmetric traffic patterns in Prefill-Decode (PD) Disaggregated Inference environments, causing hotspots and PFC backpressure at certain leaf switches.

ZCube fundamentally resolved this congestion issue by completely eliminating the spine layer and adopting a flattened structure that uses a Complete Bipartite Interconnect between two switch groups.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.