
JR
Jo\~ao Renato Ribeiro Manesco, Danilo Samuel Jodas, Douglas Rodrigues, Leandro Aparecido Passos, Jo\~ao Paulo Papa
· 1 min read
ResearcharXiv cs.CV
SPACE: Semantic Projection and Alignment of CLIP Embeddings for Domain Adaptation
arXiv:2609.23248v1 Announce Type: new
Abstract: A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to degraded performance. This challenge is amplified when the same semantic concept appears under distinct visual forms, such as photographs and sketches, where visual similarity is weak despite semantic correspondence. Existing unsupervised domain-adaptation methods aim to align distributions across domains but often ignore semantic relationships among samples of the same class. To address this issue, this paper introduces SPACE, a method that exploits the semantic structure of CLIP's vision-language space for domain adaptation. The key idea is to use text descriptions as semantic anchors by applying Singular Value Decomposition to CLIP embeddings of class descriptions, yielding an orthogonal basis that captures semantic relationships among categories. Visual features from both domains are projected into this semantic subspace, aligning images based on meaning rather than appearance.
Original source
This story was published by arXiv cs.CV and written by Jo\~ao Renato Ribeiro Manesco, Danilo Samuel Jodas, Douglas Rodrigues, Leandro Aparecido Passos, Jo\~ao Paulo Papa. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


