AI Briefing
KO

Introducing HealthBench

·2025.05.12 19:30

Key point

OpenAI has released a new benchmark, **HealthBench**, to precisely evaluate the performance and safety of medical AI.

Details

OpenAI announced a new medical AI evaluation benchmark, HealthBench, emphasizing that ensuring the usefulness and safety of models is essential for AI to contribute to improving human health.

HealthBench was built in collaboration with 262 physicians from 60 countries around the world, and includes 5,000 realistic health-related conversation datasets that reflect actual medical settings. This data supports multilingual and multi-turn conversation formats, spanning various medical specialties and personas.

This benchmark was designed based on the following three core principles:

  • Meaningful: Goes beyond simple test questions to reflect real clinical workflows and complex scenarios.
  • Trustworthy: Enhanced reliability by reflecting physicians' judgment criteria and priorities.
  • Unsaturated: Leaves room for the latest models to improve, encouraging continuous performance enhancement.

Each conversation is scored according to 48,562 unique rubrics written directly by physicians. Model responses are verified for whether each criterion is met through an evaluation model based on GPT-4.1, and the final score is calculated based on the points earned out of the total score.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.