LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
Apple
Researchers introduced LVSum, a benchmark with 72 annotated videos averaging 16 minutes to evaluate how well multimodal large language models summarize long-form video content with temporal grounding. Testing showed that transcripts significantly outperform visual frames for summarization quality, with a substantial gap remaining between AI-generated and human summaries. Current MLLMs struggle with temporal grounding and cross-modal coherence, revealing systematic weaknesses in understanding the temporal sequencing of events in extended videos.
Why it matters
Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum comprises 72 diverse videos spanning 13 domains with an average duration of 16 minutes, each annotated with up to 10 human-generated summaries containing temporal references. We conduct a comprehensive evaluation…
Related stories
Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
Apple Machine Learning Research · 1 week ago ·
17
Frame selection is the whole game: notes from making LLMs watch video
leoaido.com · 1 month ago ·
26
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
Apple · 2 months ago ·
24