
QZ
Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu
· 1 min read
ResearcharXiv cs.CV
S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection
arXiv:2609.27413v1 Announce Type: new
Abstract: Alignment-free RGB-T salient object detection (RGB-T SOD) aims to identify salient objects from unregistered RGB and thermal image pairs without costly pre-alignment. However, spatial misalignment breaks pixel-wise correspondence and causes feature contamination during cross-modal fusion. To address this issue, we propose S2A, a semantic-to-spatial alignment framework for alignment-free RGB-T SOD. Specifically, a global-guided hierarchical fusion module (GGHF) first exploits global semantic guidance to suppress background interference and refine hierarchical intra-modal features. Subsequently, the alignment-free cross-modal channel attention module (AFCA) globally exchanges complementary semantic information through channel-wise interaction, effectively overcoming the interference caused by local spatial misalignments. Finally, a spatial deformable cross-attention module (SDCA) predicts adaptive sampling offsets to recover local cross-modal spatial correspondence. Through this semantic-to-spatial paradigm, S2A first enables reliable cross-modal semantic interaction and subsequently performs local spatial calibration, effectively reducing misalignment-induced feature contamination. Without bells and whistles, S2A achieves highly competitive performance on multiple public alignment-free RGB-T benchmarks, demonstrating its effectiveness in alleviating misalignment-induced feature contamination.
Original source
This story was published by arXiv cs.CV and written by Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


