AI Briefing
KO

2x 5060 Ti Qwen3.6 Benchmark

·2026.04.28 04:37

Key point

Qwen3.6 27B and 35B-A3B were compared across multiple engines on 2x 5060 Ti 16GB.

Details

Qwen3.6-27B and Qwen3.6-35B-A3B were measured under multiple configurations on a 2-card RTX 5060 Ti 16GB setup.

  • The measurement tool was llama-benchy 0.3.5, using the conditions --pp 4096 --tg 128 --depth 0 --runs 3 --latency-mode generation --no-cache.
  • For Qwen3.6-27B, vLLM's NVFP4-MTP, TP2-PP1, no spec combination had the highest PP throughput at PP 1963 t/s, TG 38.4 t/s, TTFT 2182ms.
  • For the same model, vLLM's Lorbus/Intel AutoRound series was generally around PP 1044~1088 t/s, TG 40.2~46.9 t/s, TTFT 3792~4008ms.
  • ik-llama.cpp showed PP 1450 t/s, TG 28.4 t/s, TTFT 2945ms with IQ4_XS layer + q8_0 KV, while tensor + f16 KV showed lower PP and higher TTFT.
  • For Qwen3.6-35B-A3B, vLLM's NVFP4, TP2-PP1, no spec combination was best at PP 6259 t/s, TG 116.5 t/s, TTFT 753ms.
  • On the other hand, vLLM's DFlash n=15 dropped to TG 38.9 t/s, and ik-llama.cpp recorded TG 108.9 t/s with Unsloth Q4_K_XL layer + q8_0 KV.

The author stated that speculative decoding effectively failed due to PCIe bandwidth limitations, and mentioned plans to re-measure with larger pp/tg settings.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.