Qwen3.6 35B on a 780M iGPU
Key point
Qwen3.6 35B-A3B showed a prompt processing speed of 282 tokens per second on a Radeon 780M iGPU.
Details
On a ThinkPad T14 Gen 5 (AMD 8840U, Radeon 780M, 64GB DDR5 5600 MT/s), Qwen3.6 35B-A3B GGUF was tested using llama.cpp's Vulkan backend.
The llama-bench configuration was as follows.
-hf AesSedai/Qwen3.6-35B-A3B-GGUF:Q6_K-fa 1-ub 1024-b 1024-p 1024 -n 128 -mmp 0
The benchmark results were pp1024 282.40 t/s and tg128 20.74 t/s. According to the output, Vulkan detected the AMD Radeon 780M Graphics (RADV PHOENIX), and it ran with 99 layers offloaded.
The author noted that running at Q6 required increasing GTT and adjusting the hang timeout, and stated that it also worked normally with full context.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.