AI Briefing
KO

Zai Optimizes GLM-5.1 Inference Network

·2026.05.28 22:09

Key point

Zai cut GLM-5.1 inference infrastructure costs by 33% and boosted throughput by 15% by introducing the ZCube architecture.

Details

Zai deployed the ZCube architecture, co-developed with Tsinghua University and HarnetsAI, on a 1,000-GPU cluster, significantly improving GLM-5.1 inference performance.

Replacing the existing ROFT setup with ZCube, under the same GPU and software stack environment, achieved the following results:

  • 33% reduction in switch and optical module costs
  • 15% improvement in GPU inference throughput
  • 40.6% reduction in first-token P99 tail latency

The core of this improvement is resolving the asymmetric traffic problem that occurs during Prefill-Decode (PD) Disaggregated Inference. While the existing ROFT topology is suitable for training workloads, during PD disaggregated inference, the traffic pattern does not match the static rail mapping, causing bottlenecks and PFC backpressure at specific leaf switches.

ZCube completely eliminates the spine layer and adopts a flattened structure using a Complete Bipartite Interconnect between two switch groups. This fundamentally resolves the network congestion problem that was unavoidable under the ROFT design, achieving both cost reduction and performance improvement simultaneously.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.