LEADBOARD Benchmark for Drug Prediction Released
Key point
LEADBOARD, a benchmark applying temporal splits and noise floors to honestly evaluate drug property prediction tools, has been released.
Details
The LEADBOARD benchmark for measuring the performance of drug property prediction tools has been released. This benchmark consists of 21 boards and 18,382 holdout compounds, with labels not distributed.
1. Data Splitting Strategy Determines Performance Comparing temporal splits (training on data before 2022, testing on data after) with random splits on the hERG board resulted in an AUROC difference of 0.211. Random splits cause overestimation by mixing structurally similar molecules into training/test sets, effectively evaluating twins of what the model has already seen. In contrast, temporal splits honestly reproduce real-world unknown compound prediction scenarios and prevent overfitting because the answers do not exist at the time of training.
2. Explicit Experimental Noise Floor We calculated the noise floor by analyzing how experimental values for the same compound vary across papers. For hERG, the noise floor is 0.421 log, meaning that if the difference between 1st and 2nd place is 0.02, that ranking is statistically insignificant. This allows for an objective assessment of score reliability.
3. Baseline Performance and Limitations We ran no-learning baselines (Constant, Nearest Neighbor, Morgan+LightGBM) on all boards to reveal the limits of current technology. On some boards (CYP2D6, CYP3A4), the baseline MAE approached or fell below the noise floor, suggesting that further model improvements on these endpoints may not hold more significance than experimental inconsistency.
4. Board Tagging System Splitting methods (T: temporal, S: scaffold) and data source transparency (P1~P4) are explicitly tagged to clearly distinguish the experimental conditions of each board. Boards with random splits (R) are not opened.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.