Arena AI Model ELO History
Key point
Introduces a methodology for tracking ELO score changes and performance degradation trends of major AI models based on LMSYS Arena data.
Details
Tracks hidden trends such as performance degradation (nerf)—including excessive censorship or quantization for cost reduction—that occur during AI model updates.
Difference Between Web UI and API LMSYS Arena tests a model's original performance through API endpoints. In contrast, web interfaces used by general users (ChatGPT, Gemini, etc.) may apply system prompts, safety filters, or quantized models for cost savings during peak hours, which can create differences from API benchmarks.
Data and Analysis Logic
- Data Source: Automatically collects the LMSYS Arena Leaderboard Dataset from Hugging Face daily, based on thousands of blind human evaluations.
- Flagship Model Criteria: Among the multiple models from each AI lab, only the single flagship model with the highest ELO is tracked to generate the curve.
- Variant Model Consolidation: Reasoning mode variants such as
-thinking,-reasoningare treated as the same model and their data consolidated to prevent score fluctuations. - Trend Identification: Clearly visualizes the score increase at the time of a new model release and the degradation over the model's lifecycle.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.