AI Briefing
KO

Ground truth is a process, not a dataset

·2026.06.04 00:56

Key point

Amazon raised evaluation accuracy to 90.9% through an 'audit-then-score' protocol in which AI models challenge benchmarks and humans verify the challenges.

Details

Existing AI performance measurement methods evaluate models using fixed datasets labeled by human experts as ground truth. However, in tasks such as verifying complex research reports, even PhD-level experts only achieve an accuracy of 60.8%, clearly exposing the limits of the existing approach.

When an AI model produces a result that differs from the benchmark, rather than simply dismissing it as a model error, attention should be paid to the possibility that the benchmark itself is incomplete or incorrect. To address this, Amazon's AGI group proposed the audit-then-score protocol.

The audit-then-score protocol works as follows:

  • When an AI fact-checker disagrees with an existing benchmark, it challenges the benchmark by submitting concrete evidence and arguments.
  • A human expert (Auditor) directly compares the new evidence presented by the model with the existing benchmark's argument.
  • If the model's argument is more valid, the benchmark is revised and the model is re-evaluated.

This protocol includes a shared test set called DeepFact-Bench and the DeepFact-Eval system, which checks whether claims are supported by literature. Applying this approach substantially improved benchmark accuracy from the existing 60.8% to 90.9%.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.