AWS and Johns Hopkins Engineering Release Large-Scale Dataset for AI/ML Antibody Design
Key point
AWS and JHU have released a large-scale benchmark for evaluating AI antibody design.
Details
AWS and Johns Hopkins Engineering have released the Antibody Developability Benchmark for evaluating AI/ML-based antibody design. It is a large-scale, high-diversity benchmark designed to address the bias and limited diversity of existing public data.
Antibody development remains costly and time-consuming. Public data is often skewed toward a single format or a single target, making it difficult to reliably compare whether models actually predict developability well.
The benchmark is 20x more diverse than existing benchmarks in the literature, and includes the following composition.
- 50 seed antibodies
- 4 structural formats: IgG, VHH, NearGermline-IgG, scFv
- 42 antigens
- 6 core properties: expression, purity, thermostability, aggregation, polyreactivity, hydrophobicity
Each seed antibody was subjected to a systematic mutation strategy. It included both pLM-guided and non-pLM-guided variations, as well as amino acid substitution and insertion/deletion, producing up to 99 engineered variants per seed.
Notably, it includes both good and bad developability, and all data was validated through wet-lab experiments. Models were evaluated via zero-shot inference without prior exposure to the dataset, enabling fair comparison without data leakage.
The research team believes this benchmark will more accurately distinguish model performance and serve as an expandable benchmark to which more models and properties will be added in the future. Results are currently available on Amazon Bio Discovery, and additional benchmarks will be released along with a paper later this year.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.