AI Briefing
KO

Turning on MTP boosts throughput by 24%

·2026.04.18 21:38

Key point

On DGX Spark, Qwen3.6 performs similarly to Qwen3.5, but MTP significantly increases throughput.

Details

Comparing Qwen3.6 and Qwen3.5 on DGX Spark (GB10) with the same settings and the same benchmarks, throughput was roughly the same, within ±1% overall.

However, after turning on MTP, results changed in the 16-concurrent stress test. Throughput increased by +24%, and TTFT decreased by -57%.

The interpretation is as follows.

  • In GB10's unified memory environment, serving is close to a memory bandwidth bottleneck
  • The remaining compute headroom gets filled by MTP verification
  • This means speculative decoding delivers gains in the way it was originally intended to

Additionally, a global acceptance rate of 72.5% was presented, along with a full breakdown by workload type.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.