NPHardEval Leaderboard for LLM Reasoning Evaluation Released
·2024.02.02 09:00
Key point
The NPHardEval leaderboard, which dynamically evaluates LLMs' logical reasoning ability using complexity classes, has been released.
Details
NPHardEval is a benchmark developed by researchers at the University of Michigan and Rutgers University that quantitatively measures LLM reasoning ability using Complexity Classes.
Key Features
- Dynamic Updates: Questions are automatically generated and updated monthly to prevent model Overfitting.
- Logic-focused Evaluation: Numerical computation problems are excluded to reduce interference from numerical operations, focusing on evaluating pure logical reasoning ability.
- Use of Complexity Hierarchy: Algorithmic problems spanning P, NP-complete, and NP-hard are used to systematically classify the depth of reasoning.
Data and Metrics
- A total of 900 questions are provided by combining 9 algorithms with 10 levels of difficulty.
- Weighted Accuracy, which applies weights by difficulty level, and Failure Rate, which measures the appropriateness of output format, are used as the main metrics.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.