AI Briefing
KO

Qwen 2x Speedup

·2026.04.15 11:46

Key point

With DFlash in oMLX 0.3.5 RC1, Qwen3.5 27B went from 9 to 22 T/S on a Mac M5 Max.

Details

This is an early test result showing that enabling DFlash support in oMLX 0.3.5 RC1 boosted Qwen3.5 27B (BF16) generation speed from 9 T/S to 22 T/S on a Mac M5 Max 128GB.

  • Main model: Jackrong/MLX-Qwopus3.5-27B-v3-bf16
  • Draft model: z-lab/Qwen3.5-27B-DFlash
  • Environment: M5 Max 128GB
  • The author shared links to the DFlash repository and oMLX RC1 together, highlighting the potential for noticeable speed improvements in local deployment.

The author noted that since Qwen3.5 27B already offers good performance relative to its size, solving the speed issue makes it much more practical to use with higher quants or full weights.

However, it hasn't been tested yet in other harnesses like OpenCode, and these numbers are still early-stage results.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.