
ES
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
· 1 min read
ResearcharXiv cs.CV
Video-Index: A Curated Meta-Benchmark for Video Understanding
arXiv:2610.00960v1 Announce Type: new
Abstract: A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index
Original source
This story was published by arXiv cs.CV and written by Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


