AI Briefing
KO

OpenAI o1 correctly diagnosed 67% of ER patients while triage doctors scored 50-55%

·2026.05.04 10:34

Key point

OpenAI o1 outperformed doctors in emergency room note-based diagnosis and treatment planning.

Details

In a Harvard study published in Science on April 30, 2026, OpenAI o1 was applied to initial ER triage and clinical reasoning tasks and compared against human doctors.

  • The model was given only standard electronic medical records for 76 Boston ER patients, with records including vital signs, demographic information, and nurses' visit-reason notes.
  • o1 provided a correct or very close diagnosis 67% of the time, while two human doctors scored 50-55%.
  • When given more information, o1 scored 82% and expert doctors scored 70-79%, but this difference was not statistically significant.
  • For long-term treatment planning such as antibiotic regimens and end-of-life care planning, o1 also outperformed 46 doctors, scoring 89% vs. 34% across 5 clinical cases.
  • The comparison was limited to information that could be conveyed as text, and did not include patients' appearance, distress, or non-verbal cues.

A separate survey found that about 20% of US doctors use AI for diagnostic assistance, while in the UK 16% use it daily and 15% use it weekly. Errors and accountability remain unresolved issues, the report noted.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.