SyncAI.news, a Varaisys broadcasting
Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
AF

Anas Filali Razzouki, Killian Steunou, Khalil Guetari, Thomas Kling, Moun\^im El-Yacoubi, Yannis Tevissen

· 1 min read

ResearcharXiv cs.CV

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

arXiv:2610.10163v1 Announce Type: new Abstract: Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manually verified and filtered to create a benchmark of 750 captioned clips and 3,000 person-centric questions. We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models. Combining visual face boxes with textual coordinates yields the most consistent performance across scales and significantly improves overall performance over coordinates alone. Smaller models tend to over-assign known identities when the queried person is not grounded, while larger models better recognize such UNIDENTIFIED cases. We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips. Across all scales, BAC outperforms every other evaluated model family of comparable size. BAC-8B reaches 93.20\% overall QA accuracy, ranking behind only GPT-5.6 Sol among the frontier models evaluated in our study. Overall, explicitly communicating who is where, together with lightweight task-specific adaptation, substantially improves identity-aware video understanding without changing the underlying architecture. We release the benchmark, training data, code, and BAC checkpoints at https://github.com/momentslab/beyond-anonymous-captions.

Original source

This story was published by arXiv cs.CV and written by Anas Filali Razzouki, Killian Steunou, Khalil Guetari, Thomas Kling, Moun\^im El-Yacoubi, Yannis Tevissen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News