ADeLe: Predicting and Explaining AI Performance Across Diverse Tasks
Key point
It explains LLM performance with 18 ability scores and predicts new task success rates with 88% accuracy.
Details
Existing AI benchmarks only show how well a model does on a particular task, but don't explain well why it succeeds or fails. ADeLe (AI Evaluation with Demand Levels) scores both models and tasks against the same 18 core abilities, allowing performance to be read as an ability structure rather than a simple total score.
Each task is assigned a demand level from 0 to 5 for abilities such as attention, reasoning, and domain knowledge, and each model builds an ability profile based on its performance across multiple tasks. By comparing the resulting model profile with the task demand profile, it becomes possible to pinpoint exactly which ability gap caused a failure.
Using this approach, the team examined what various benchmarks actually measure. The results revealed that many benchmarks fail to cleanly isolate the abilities they intend to measure, or have too narrow a difficulty range, giving an incomplete picture of model ability.
The team also evaluated 15 LLMs to build ability profiles, using the difficulty level corresponding to a 50% success probability for each ability as that ability's score. As a result, they predicted new task success or failure for models like GPT-4o and LLaMA-3.1-405B with about 88% accuracy, outperforming existing methods.
In particular, reasoning benchmarks that appeared to belong to the same category actually required significantly different abilities. Some tasks involved only basic problem-solving, while others required advanced logic, abstraction, and domain knowledge together, and the same model could score above 90% on low-difficulty tasks but drop below 15% on high-difficulty ones.
Microsoft researchers explain that reasoning-oriented models like o1 and GPT-5 show improved performance in logic, math, and interpreting user intent, but performance drops as task demand increases. ADeLe quantifies these limitations, showing within a single framework how far a model can reason and where it breaks down.
Going forward, this could be extended to multimodal and embodied AI, and used as a standard evaluation framework for AI research, policymaking, and security auditing. The key is evolving beyond simple score comparisons into an evaluation method that simultaneously provides performance prediction and error explanation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.