Hugging Face Supports AWS Inferentia2
·2024.05.22 09:00
Key point
Hugging Face now supports AWS Inferentia2, enabling deployment of over 100,000 models on SageMaker and Inference Endpoints.
1 / 2
Details
Hugging Face has expanded its ability to deploy models via AWS Inferentia2. Users can now efficiently serve models using Inferentia2 instances on Hugging Face Inference Endpoints and Amazon SageMaker.
Key Update Details:
- Expanded SageMaker Support: Over 100,000 public models, including Llama 3, across 14 model architectures such as BERT, RoBERTa, and ViT, and 6 tasks including text classification and question answering, can be deployed on Inferentia2.
- Introduction of Inference Endpoints: Models can be deployed by selecting Inferentia2 instances in just a few clicks.
- Inf2-small: 2 cores, 32GB memory (suitable for Llama 3 8B, $0.75/hour)
- Inf2-xlarge: 24 cores, 384GB memory (suitable for Llama 3 70B, $12/hour)
- Leveraging TGI (Text Generation Inference): When deploying LLMs like Llama 3, TGI for Neuron is used to support continuous batching and streaming features, and it is compatible with the OpenAI SDK, allowing use without changes to existing application code.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.