AI Briefing
KO

R9700 Ngram-Mod Speed Test

·2026.04.27 13:12

Key point

Applying ngram-mod to Qwen3.6 27B on the R9700 showed speed improvements, though with some variability.

Details

Tested Qwen3.6 27B on an AMD Radeon AI PRO R9700 using llama.cpp's --spec-type ngram-mod.

  • With llama-bench-vulkan, recorded pp512 1050.13 t/s and tg128 31.26 t/s.
  • llama-server-vulkan was run with settings such as --spec-ngram-size-n 24, --draft-min 12, --draft-max 48.
  • Token generation speed showed average 28.80 t/s, median 28.20 t/s, P95 45.34 t/s.
  • Prompt processing speed showed average 549.60 t/s, median 519.19 t/s, P95 936.60 t/s.
  • The author concluded that while there is variability in performance, the speed increase was useful for work involving the same codebase.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.