Hugging Face Launches Inference API for PRO
Key point
Hugging Face has launched 'Inference for PRO,' offering PRO subscribers a dedicated API for high-performance models with higher usage limits.
Details
Hugging Face has launched Inference for PRO, a service providing high-speed inference APIs for the latest models to PRO subscribers.
This service leverages text-generation-inference to provide dedicated HTTP endpoints capable of ultra-fast inference for a curated selection of models. PRO users receive a higher Rate Limit than the existing free Inference API, and can immediately use the following key models:
- LLM: Meta Llama 3 (8B, 70B), Mixtral 8x7B, Zephyr 7B, Code Llama, and more
- Image Generation: Stable Diffusion XL
- Audio Generation: Bark
This service is optimized for quickly testing and prototyping the latest models without building your own infrastructure. However, for building large-scale production environments, using the dedicated Inference Endpoints service is recommended.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.