
ZZ
Zhengtong Zhu, Jiaqing Fan, Hanwen Qian, Fanzhang Li
· 1 min read
ResearcharXiv cs.CV
SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation
arXiv:2609.31725v1 Announce Type: new
Abstract: Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the performance of the model in sparse video scenes, while existing VOS methods have limited adaptability to global changes in objects. Based on this observation, in this paper, we propose self-supervised VOS with Sliding window, Wavelet transform and optimal Transport (SWT), a self-supervised VOS framework entirely trained on static dataset using contrastive learning. Firstly, a rolling sample buffer reuses overlapping groups of independently sampled images across successive updates. Secondly, to address the long-distance modeling difficulty caused by simple convolutional structures, we introduce wavelet transform to expand the receptive field of convolutional kernels, thus improving the model's representational capability. Finally, we incorporate optimal transport to assist the model in finding the globally optimal match between the target across two frames, improving the model's ability to handle nonrigid deformations of objects. SWT only requires training on the COCO dataset once and achieves excellent results on five VOS datasets as well as an additional body part propagation dataset. The code will be released soon at [https://github.com/machine928/SWT.git](https://github.com/machine928/SWT.git).
Original source
This story was published by arXiv cs.CV and written by Zhengtong Zhu, Jiaqing Fan, Hanwen Qian, Fanzhang Li. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


