ONNX Runtime Accelerates 130,000 Models
Key point
ONNX Runtime supports over 130,000 models on Hugging Face, accelerating inference performance.
Details
ONNX Runtime is a cross-platform machine learning tool that can optimize the performance of various models.
Hugging Face currently has over 130,000 ONNX-supported models, which can significantly improve models' inference performance. For example, running the Whisper-tiny model with ONNX Runtime can reduce average inference latency by up to 74.30% compared to PyTorch.
Currently, over 90 architectures are supported, including the following major model architectures:
- NLP models such as BERT, GPT2, DistilBERT, RoBERTa, T5
- Image generation models such as Stable-Diffusion
- Audio models such as Whisper, Wav2Vec2
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.