Hugging Face Adds Google TPU Support
Key point
Google Cloud TPU v5e is now available in Hugging Face's Inference Endpoints and Spaces.
Details
Through a collaboration with Google Cloud, Hugging Face now supports Google Cloud TPU v5e in Inference Endpoints and Spaces. AI developers can use this to achieve higher performance and cost efficiency when deploying models and running demos.
The TPU v5e configurations available in Inference Endpoints are as follows:
- v5litepod-1: 1 core, 16GB memory ($1.375/hour)
- v5litepod-4: 4 cores, 64GB memory ($5.50/hour)
- v5litepod-8: 8 cores, 128GB memory ($11.00/hour)
For smooth operation of large-scale models, a configuration of v5litepod-4 or higher is recommended.
Hugging Face Spaces also supports the same three TPU v5e instance configurations, enabling quick building of AI demos.
As part of this collaboration, the open-source library Optimum TPU was developed, making it easy to deploy major models such as Gemma, Llama, and Mistral.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.