AI Briefing
KO

TGI Begins Native Support for Intel Gaudi

·2025.03.28 09:00

Key point

HuggingFace's LLM serving solution TGI now natively supports Intel Gaudi hardware.

Details

Intel Gaudi hardware support has been natively integrated into HuggingFace's LLM serving solution, Text Generation Inference (TGI). Previously supported through a separate fork, it is now supported directly in the main codebase via TGI's new multi-backend architecture.

Supported hardware includes the full Gaudi1, Gaudi2, Gaudi3 lineup, and it is available through AWS, Intel Tiber AI Cloud, IBM Cloud, and major OEMs (Dell, HP, Supermicro).

Key Features and Benefits:

  • Hardware Diversity and Cost Efficiency: Offers deployment options beyond GPUs, providing high cost-performance for specific workloads.
  • Production Environment Optimization: Supports dynamic batching, streaming responses, multi-card inference (sharding), and FP8 precision.
  • Model Support: Provides optimized performance for major models such as Llama 3.1/3.3, Mistral, Mixtral, Qwen2, Gemma.

There are plans to expand support to include the latest models such as DeepSeek-r1/v3 in the future.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.