AI Briefing
KO

SimpleQA Released

·2024.10.30 19:00

Key point

OpenAI has open-sourced SimpleQA, a new benchmark for measuring the factuality of language models.

Details

To measure and reduce hallucinations, a persistent problem in language models, OpenAI has open-sourced the SimpleQA benchmark, which evaluates the ability to answer short, fact-seeking questions.

SimpleQA has the following characteristics:

  • High accuracy: Uses answers verified by two independent AI trainers.
  • Diversity: Covers a wide range of topics, from science and technology to TV shows and video games.
  • High difficulty: More challenging than existing benchmarks, to the point that GPT-4o scores below 40%.
  • Great research UX: Questions and answers are concise, enabling fast execution and efficient grading.

For the quality of the dataset, questions must have a single, indisputable answer, and the facts must not change over time. Verification estimates the dataset's inherent error rate at about 3%.

Grading uses a ChatGPT classifier, which categorizes model answers into three grades: 'Correct,' 'Incorrect,' and 'Not attempted.'

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.