Qwen-3.6-27B Speed Experiment
Key point
With llama.cpp ngram-mod speculative decoding, Qwen-3.6-27B speed rose to 136.75 t/s.
Details
In llama.cpp llama-server, Qwen3.6-27B-Q8_0.gguf and mmproj-BF16Qwen3.6-27B.gguf were attached, and --spec-type ngram-mod --spec-ngram-size-n 24 --draft-min 12 --draft-max 48 was applied.
During the session, generation speed climbed 13.60 t/s → 25.53 t/s → 68.35 t/s → 136.75 t/s, and at each stage Qwen generated the full code to completion.
When shown a screenshot with the browser console open to demonstrate a bug, the model pinpointed the cause and fixed it, and in the end the completeness of the aquarium example improved significantly.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.