M5pro ran 122B
·2026.04.14 03:56
Key point
On M5pro 64GB, Qwen 3.5 122B ran fairly well at IQ2-m.
Details
Testing Qwen 3.5 122B on M5pro 64GB, IQ2-m worked quite well while IQ3-xxs was close to the memory limit.
- Total load memory at IQ3-xxs: 50.83GB
- Model's own usage: 44.76GB
- Total system RAM usage at 32k ctx: 59.07GB / 64GB
- GPU offload: 48/48
- Lowering to 46/48 made almost no perceptible difference, only reducing token speed by about 2 tok/s
- Output speed: 36 tok/s at the start, 28 tok/s upon reaching 32k ctx
- Time to first token: 6-8 seconds across the board
Even a large model could be run locally on an M5pro 64GB environment with sufficient unified memory, and in particular IQ2-m looks meaningfully viable for practical use.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.