AI Briefing
KO

Luce DFlash Support for AMD Strix Halo

·2026.05.13 03:09

Key point

Luce DFlash/PFlash boosts LLM inference performance by up to more than 3x on the AMD Strix Halo iGPU.

Details

The open-source project lucebox-hub has released DFlash and PFlash support for the AMD Ryzen AI MAX+ 395 (Strix Halo) iGPU. This is the result of optimizing the existing Luce DFlash stack for consumer AMD APU environments.

Benchmark results comparing performance against llama.cpp HIP on the Qwen3.6-27B (Q4_K_M) model are as follows:

  • Decode: 26.85 tok/s (2.23x improvement)
  • Prefill (16K context): 20.2s (3.05x improvement)

For a workload of a 16K prompt and 1K generation, total execution time was reduced from 147 seconds to 58 seconds, recording 2.5x faster end-to-end performance.

In particular, by leveraging 128 GiB of unified memory, it offers the advantage of efficiently running large-scale models (such as Qwen3.5-122B-A10B) locally that are difficult to run on consumer GPUs with the existing 24 GiB VRAM.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.