TLDRocket
Sign in

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

Apple ML Research

Apple built a new test to see if AI can actually summarize long videos with correct timestamps, not just describe scenes. Turns out top AI models still fumble timing and lean on transcripts more than the actual video.

Apple's ML research team just dropped LVSum, a benchmark designed to expose a weakness that's been quietly ignored in video AI: models are bad at knowing when things happen. Not just what happens in a video, but when — and that distinction matters a lot if you want an AI to summarize a 16-minute clip and actually point to the right moment.

The setup is deliberately grueling. LVSum has 72 videos across 13 different domains, averaging 16 minutes each, and every single one comes with up to 10 human-written summaries that include actual timestamped references. That's a lot of manual annotation work, and it's the kind of dataset that doesn't get built unless someone genuinely wants to stress-test whether multimodal large language models can hold onto a timeline over an extended stretch rather than just describing isolated frames.

The team ran a batch of leading proprietary and open-source MLLMs through this gauntlet, using both standard automatic scoring and new LLM-based metrics they built specifically to judge relevance and how well the video and text modalities line up. Three things fell out of the results. First, transcripts do most of the heavy lifting — models lean on spoken or written text far more than on visual frames when producing decent summaries, which says something about how shallow current visual temporal understanding really is. Second, even the best model output still trails noticeably behind what a human writes, in both content and structure. Third, and probably the most damning finding, is that these models keep messing up temporal grounding specifically — they struggle to follow instructions about timing and often produce summaries where the video and text don't cohere the way they should.

None of this is shocking if you've watched how MLLMs get trained. Most training objectives reward getting the gist of a scene right, not tracking how events unfold and connect across minutes of footage. LVSum essentially puts a number on that gap, and it's a bigger gap than the marketing decks for these models would suggest.

My take

This is the kind of benchmark that quietly matters more than another flashy demo — it shows the industry has been grading video AI on vibes, not on whether it actually understands time. I'd bet real money that most 'video understanding' claims from labs this year would collapse under LVSum's timestamp scrutiny, and that's exactly why independent, annoying benchmarks like this need to exist.

Read more about this at: Apple ML Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.