
MS
Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski
· 1 min read
ResearcharXiv cs.CV
Scalable Patch-Level Self-Supervised Learning
arXiv:2610.10013v1 Announce Type: new
Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on $12\times$ less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.
Original source
This story was published by arXiv cs.CV and written by Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


