AI Briefing
KO

122B 198t/s

·2026.04.10 09:59

Key point

Verified Qwen3.5-122B running at 198 tok/s on 2x RTX PRO 6000 Blackwell.

Details

The author shared results running Qwen3.5-122B NVFP4 at about 198 tok/s using a combination of 2x RTX PRO 6000 Blackwell 96GB, EPYC 4564P (AM5), 128GB DDR5 ECC, and a PCIe switch.

The key point is not the speed itself but the cost ratio. The author concluded that this build can reach the same 2-GPU inference performance as a Threadripper Pro workstation, but at a cheaper cost. However, the original claim that "switch topology is faster" was revised, and it was noted that both switch and direct-attach show the same measured P2P latency of 0.38 µs.

The benchmarks are as follows.

  • Qwen3.5-122B NVFP4: about 198 tok/s
  • Qwen3.5-27B FP8: 169.7 tok/s
  • MiniMax M2.5 NVFP4: 148.1 tok/s
  • Qwen3.5-122B NVFP4 (vLLM): 131.4 tok/s
  • Qwen3.5-397B GGUF: 79 tok/s

The 122B result is the average of 3 measurements of 200.3 / 206.7 / 190.2 tok/s, and the variability was explained as being due to FlashInfer autotuner non-determinism.

Another key point is that on direct-attach rigs, when the topology is NODE/PHB, the ForceP2P module setting may be required. Without this setting, it was noted that P2P writes fall through to SysMem staging, and --enable-pcie-oneshot-allreduce effectively falls back to NCCL, which can degrade performance.

In summary, the conclusion of this post is a hands-on report that with a relatively cheaper platform like AM5 EPYC + c-payne PM50100, using the combination of SGLang b12x + NEXTN speculative decoding can achieve Threadripper Pro-level inference performance on 2x RTX PRO 6000 Blackwell.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.