AI Briefing
KO

vLLM Adds HIP W4A16 Kernel Support for AMD

·2026.05.29 21:31

Key point

A native HIP W4A16 kernel that optimizes performance on AMD GPUs has been added to vLLM.

Details

A native HIP W4A16 kernel that optimizes inference performance for 4-bit quantized models on AMD GPU (ROCm) environments has been merged into the vLLM project.

With this update, inference speed for W4A16 quantized models on RDNA3-architecture GPUs has significantly improved. Key benchmark results based on the Qwen3.6-27B-GPTQ-W4A16-G32 model are as follows:

  • RDNA3 W4A16 (new kernel) bf16: 205.3 tk/s
  • RDNA3 W4A16 (new kernel) fp16: 270.2 tk/s

With improved performance over the existing Triton approach, this is expected to offer greater efficiency for developers looking to serve LLMs using AMD hardware.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.