AI Briefing
KO

M2 MacBook Pro 32GB Qwen 27B

·2026.04.29 06:01

Key point

On an M2 MacBook Pro with 32GB RAM, Qwen 3.6 27B's processing speed plummeted to 4t/s at a 52,000-token context.

Details

Ran Qwen 3.6 27B on an M2 MacBook Pro with 32GB RAM using 4bit XS unsloth quant.

Used the model files Qwen3.6-27B-IQ4_XS.gguf and Qwen3.6-27B-mmproj-BF16.gguf.

Initially it was prompt processing 80 t/s and generation 7.9 t/s, but once the context reached 52,000 tokens, it dropped to 4 t/s pp and 3.1 t/s tg.

  • No signs of swapping were observed, and memory pressure never exceeded the yellow level.
  • llama-server was run with the settings -c 131072, -batch-size 256, -ngl 99, -ctk q8_0, -ctv q8_0, --spec-type ngram-mod, --spec.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.