AI Briefing
KO

Apple Releases CapQuiz, a Reference-Free Benchmark for Video Caption Quality Assessment

·2026.09.11 09:00

Key point

Apple researchers have released CapQuiz, a reference-free benchmark for evaluating the information fidelity of video captions.

Details

Existing video caption evaluation metrics rely on matching generated text with ground-truth text, causing high-quality captions with different expressions or shifted perspectives to be disadvantaged. Additionally, evaluations were fragmented, making detailed analysis of caption quality difficult.

To overcome these limitations, Apple researchers introduced CapQuiz, a new reference-free benchmark. This benchmark assesses quality by measuring how well captions can answer human-verified, granular multiple-choice questions extracted from videos. The questions are classified into 10 types across the Descriptive and Inferential categories and cover 24 diverse video domains.

The researchers proposed a composite metric called CapF1 to comprehensively measure caption Factuality and Coverage. CapF1 combines CapP (Factuality) and CapR (Coverage). Experimental results confirmed that CapQuiz shows a significantly higher correlation with human judgment than existing metrics and provides interpretable insights into model performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.