Open Agent Leaderboard Unveiled
Key point
Hugging Face has released a leaderboard that compares agent systems' performance and cost together.
Details
The Open Agent Leaderboard has been released, comparing not models but entire agent systems. Alongside performance, it also discloses cost, allowing users to judge which combinations offer value in actual deployment.
The evaluation brought together tasks of different characters—SWE-Bench Verified, BrowseComp+, AppWorld, tau2-Bench Airline & Retail, tau2-Bench Telecom—under a single common protocol. Each benchmark was standardized into a task / context / actions structure, preserving existing designs while enabling direct comparison of results across agents.
The key findings are as follows.
- Even with the same model, scores and costs varied significantly depending on the agent design.
- The top 3 combinations used the same model, but results diverged due to differences in agent implementation.
- General-purpose agents performed similarly to or better than specialized systems on some tasks.
- Failed runs were 20-54% more expensive than successful runs, and
tool shortlisting, which narrows down tool candidates, boosted performance across all models.
The leaderboard has now expanded to a total of 5 models, 5 agents, and 6 benchmarks, including DeepSeek V3.2 and Kimi K2.5. Open-weight models score on average 18-29 points lower than top-tier closed models, and the Exgentic framework, along with the results dataset and paper, have been released together.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.