SyncAI.news, a Varaisys broadcasting
S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization
ZM

Ziqian Mo, Hill Zhang, Haosheng Tan, Ling Li, Jiaheng Wei

· 1 min read

ResearcharXiv cs.CV

S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

arXiv:2610.11608v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views. Existing methods mainly rely on visual representations, but often fail to jointly model fine-grained structural correspondences and semantic priors, making them prone to confusion between visually similar but semantically different regions, and thus limiting robustness under large viewpoint variations. To address these challenges, we propose \textbf{S$^3$Geo}, a structure-semantic synergistic learning framework for cross-view matching. Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns. We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment. Furthermore, we incorporate a Semantic Knowledge Distillation (SKD) strategy from a frozen CLIP teacher to transfer semantic priors and relational structures, thereby improving discrimination on hard negatives. By operating synergistically, the semantic priors provide robust contextual filtering, which guides the structural module to establish precise spatial alignments. Experiments on the University-1652 and SUES-200 datasets demonstrate that \textbf{S$^3$Geo} consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.

Original source

This story was published by arXiv cs.CV and written by Ziqian Mo, Hill Zhang, Haosheng Tan, Ling Li, Jiaheng Wei. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News