AI Briefing
KO

DTree Comes Back to Win

·2026.04.16 01:05

Key point

On an M2 Max, DTree showed MLX results **1.07x** faster than DFlash.

Details

In an M2 Max 32GB environment, running Qwen3.5-4B with q4_g64, and matching spec=16, tree_budget=24 settings, the following comparison results came out.

  • DFlash: 45.07 e2e tok/s
  • DTree: 48.31 e2e tok/s

On a local MLX basis, DTree came out about 1.07x ahead. It's not a huge gap, but since repeated tests showed similar results, it was judged to be a meaningful result.

Most other attempts were similar or slower, and the interpretation is that for now, MLX verifier cost appears to be the main bottleneck for performance improvement.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.