AI Model Outperforms Emergency Room Doctors in Diagnosis
Key point
In a Harvard/BIDMC study, an OpenAI reasoning model outperformed emergency room doctors in diagnostic performance.
Details
In a study published in Science by Harvard Medical School and Beth Israel Deaconess Medical Center, an OpenAI reasoning model outperformed 2 experienced physicians in diagnosing emergency room patients and making treatment decisions.
The research team evaluated the model using a combination of real patient cases, NEJM case reports, and clinical vignettes, comparing diagnostic accuracy at three time points from emergency room triage to hospital admission, given only electronic health records (EHR). Under the same conditions, the model also produced better results than the previous GPT-4.
- It showed meaningful performance even on noisy real-world emergency room records.
- Its differential diagnosis ability, which broadly suggests possible conditions, improved significantly.
- However, since only text was used, imaging, auscultation, and nonverbal cues were not reflected.
The research team did not view these results as grounds for replacing physicians, and emphasized the need for prospective clinical trials to verify whether incorporating this into actual clinical workflows improves patient outcomes.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.