AI Briefing
KO

Qwen 3.6 35B GGUF Released

·2026.05.21 00:42

Key point

ByteShape released Qwen 3.6 35B GGUF NTP and MTP quantizations along with benchmarks.

Details

ByteShape released Qwen 3.6 35B GGUF in two lines, NTP and MTP.

  • In NTP, the strategy of "picking the largest quant that fits" worked most stably.
  • A lower bpw was not always faster or better, and the largest deployable build often had the edge in both quality and speed.
  • MTP generally boosted GPU generation speed by 20~40%, but the extra memory usage could change which configurations were actually feasible.
  • MTP's gains were workload-dependent, so they didn't guarantee a consistent speedup in every situation.
  • On CPU environments, MTP's appeal was lower, so the CPU recommendation remained NTP.
  • MMLU was excluded from this release because it could muddy the comparison signal.

The comparison was structured less like a simple model upload and more like a small-scale hardware study, with broad measurements across RTX 4090, 5090, Pro 6000, 4080, 5060 Ti and Intel i7, Intel Ultra 7, Ryzen 9, Raspberry Pi 5.

In practical terms, it emphasizes that even for the same model, a smaller bpw is not always the right answer, and if you have memory headroom, a larger quant can be the better choice in both quality and speed.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.