SyncAI.news, a Varaisys broadcasting
TempCore: Are Video QA Benchmarks Temporally Grounded?
HO

Hyunjong Ok, Jaeho Lee

· 1 min read

ResearcharXiv cs.CV

TempCore: Are Video QA Benchmarks Temporally Grounded?

arXiv:2509.01167v3 Announce Type: replace Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current Video QA benchmarks genuinely require temporal frame selection, or can most questions be answered regardless of which frames are shown? We introduce Frame Selection Sensitivity (FSS), a per-sample diagnostic that measures how much VLM accuracy changes when the most relevant frames are replaced with the least relevant ones. Across six benchmarks and eight VLMs, we find that a large majority of samples are frame-agnostic: only a minority are genuinely sensitive to frame choice. Combining FSS with a Language Independence Score (LIS) reveals that merely 5.5--31% of samples are Temporally Sensitive. We construct TempCore, compact evaluation subsets that isolate these temporal samples from existing benchmarks, and will release code and per-sample annotations upon publication.

Original source

This story was published by arXiv cs.CV and written by Hyunjong Ok, Jaeho Lee. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News