
AA
Akash Abdu Jyothi, Greg Mori
· 1 min read
ResearcharXiv cs.CV
Video-STLayout Pre-training
arXiv:2609.24031v1 Announce Type: new
Abstract: In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.
Original source
This story was published by arXiv cs.CV and written by Akash Abdu Jyothi, Greg Mori. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


