AI Briefing
KO

Measuring Model Performance on Real-World Work Tasks

·2025.09.25 18:00

Key point

OpenAI has unveiled a new evaluation metric, **GDPval**, which measures the economic value of AI through practical tasks across 44 occupations.

Details

OpenAI has introduced a new evaluation metric called GDPval to measure AI models' ability to create real economic value. GDPval evaluates models' practical work performance across 9 industry sectors and 44 occupations that contribute significantly to the US GDP.

Going beyond academic benchmarks like MMLU or evaluations focused on specific coding tasks like SWE-Bench, GDPval is based on the actual outputs of knowledge work performed by real professionals. The evaluation items include a variety of outputs used in actual work, such as legal drafts, engineering blueprints, customer support conversations, and nursing care plans.

The key features of GDPval are as follows:

  • Realism: Rather than simple text prompts, it provides reference files and context together.
  • Diversity: It requires results in various forms, including documents, slides, diagrams, spreadsheets, and multimedia.
  • Expertise: Experts with an average of 14+ years of experience directly designed and verified 1,320 professional tasks.

Currently, GDPval is conducted as a one-shot evaluation, which has limitations in fully reflecting complex workflows where models build context through multiple revisions. Future versions are planned to expand toward more interactive workflows and richer context.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.