AI Briefing
KO

MLE-bench: A Machine Learning Agent Benchmark for Evaluating Machine Learning Engineering Capabilities

·2024.10.10 19:00

Key point

OpenAI has released MLE-bench, a benchmark for measuring the machine learning engineering capabilities of AI agents.

Details

We introduce MLE-bench to measure how well AI agents perform real-world machine learning engineering (MLE) tasks. This benchmark is built on 75 machine learning-related competitions from Kaggle, and tests practical skills such as model training, dataset preparation, and running experiments.

We established human baselines using each competition's public leaderboard. As a result of evaluating several frontier language models, the setup combining OpenAI's o1-preview with AIDE scaffolding showed the best performance, achieving at least a Kaggle bronze medal level in 16.9% of all competitions.

We also investigated how agents scale with resources and the impact of pre-training data contamination. To help future research into the ML engineering capabilities of AI agents, we have released the benchmark code as open source.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.