LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
Key point
A human-annotated LVSum benchmark has been released to evaluate temporal accuracy in long video summarization.
Details
It is a very challenging task for multimodal large language models (MLLMs) to maintain temporal fidelity and generate summaries grounded in both semantics and time when summarizing long videos.
To address this, LVSum has been introduced. LVSum includes 72 diverse videos across 13 domains with an average length of 16 minutes, and each video is accompanied by up to 10 human-written summaries that include temporal reference information.
The key experimental findings are as follows:
- Importance of text: Transcript data contributes far more to improving summary quality than visual frames
- Performance gap: A significant performance gap still exists between model-generated summaries and human-written summaries
- Limitations of MLLMs: Current MLLMs show systematic weaknesses in Temporal Grounding, instruction following, and cross-modal consistency
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.