Real-Time Model Performance Metrics Now Available via AI Gateway
Key point
Vercel AI Gateway now provides real-time throughput and latency metrics for hundreds of models.
Details
AI Gateway has begun providing real-time throughput and latency metrics for hundreds of models. Users can select the model best suited to their project based on actual performance data.
Metrics are updated hourly through the following three routes:
- Model list: Displays each model's optimal performance (P50 latency and throughput)
- Model detail pages: Provider-by-provider performance breakdown for the same model
- REST API: Provides rolling endpoint performance aggregate data (P50/P95 latency and throughput)
In the model list, you can sort by throughput or latency to easily find the model with the fastest token generation or the shortest TTFT (Time-to-First-Token). You can also compare performance across multiple providers offering the same model via the detail pages.
These metrics can also be accessed programmatically via the REST API, making real-time performance data available in automated environments as well.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.