AI Briefing
KO

Hugging Face Supports AMD MI300

·2024.05.21 09:00

Key point

Hugging Face has integrated AMD Instinct MI300 GPUs into its platform to support model deployment and inference optimization.

Details

Through collaboration between Hugging Face and AMD, AMD Instinct MI300 GPUs have been integrated into the Hugging Face platform as first-class citizens. Users can now prototype and deploy models on AMD hardware using the transformers and text-generation-inference(TGI) libraries without any code changes.

Key updates include the following:

  • Infrastructure and CI/CD: Leveraging Azure's ND MI300x V5 VMs, a continuous integration and deployment (CI/CD) environment has been built. This ensures that the transformers and TGI libraries are regularly tested on MI250 and MI300 to guarantee stability.

  • Inference Performance Optimization: In collaboration with AMD engineers, Flash Attention v2, Paged Attention, GPTQ/AWQ quantization, ROCm TunableOp, and optimized fused kernels have been integrated to maximize the inference performance of models such as the Llama family.

Through this integration, developers can now efficiently run and operate large-scale AI models using AMD accelerators, similar to the NVIDIA environment.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.