Apex-Testing Coding Benchmark Updated
Key point
Apex-Testing, which evaluates AI agents' coding abilities using real GitHub repositories, has been updated.
Details
Apex-Testing is a real-world agent coding benchmark designed to address the data bias and 'benchmaxxing' problems of existing benchmarks.
Using 65-70 real private GitHub repositories, it tests whether AI models can understand codebases and perform bug fixes or feature implementations like real developers.
Key Features and Provided Metrics:
- Practice-oriented Tasks: Consists of 70 tasks across 8 categories, reflecting real work environments.
- Various Evaluation Metrics: Provides average cost, time spent, scores by difficulty level, and an ELO-based leaderboard.
- Model Comparison: Enables objective comparison of the latest LLMs' agent coding performance.
Updates for the latest models, including Qwen and DeepSeek, are currently in progress.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.