SyncAI.news, a Varaisys broadcasting
Gauge-Equivariant Attention for Rotation-Stable $360^\circ$ Scene Understanding
TZ

Tianjian Zhou, Yishan Li, Jie Jiang, Yifei Zhang

· 1 min read

ResearcharXiv cs.CV

Gauge-Equivariant Attention for Rotation-Stable $360^\circ$ Scene Understanding

arXiv:2609.31755v1 Announce Type: new Abstract: Panoramic $360^\circ$ scene understanding increasingly relies on icosphere transformers, but a state-of-the-art spherical model loses more than half of its segmentation accuracy when the camera rotates by $90^\circ$, and controlled ablations identify gauge dependence in its relative-position bias as a major contributor. We propose gauge-equivariant relative position encoding (GE-RPE): a parameter-free Reynolds average of the bias over a finite cyclic subgroup $C_n\!\subset\!\mathrm{SO}(2)$ of gauge rotations. Plugged into a SphereUFormer backbone the change is invisible at deployment---zero added parameters and $1.5$--$4.4\%$ forward latency---and the matched three-seed GE-RPE model records a $1.3\%$ drop; the published SphereUFormer checkpoint records $53\%$ under the same stress protocol but a different training recipe. Once the gauge defect is removed and a teacher-token permutation $\pi_R$ aligns the SSL views to the rotated student frame, iBOT$+$MAE pretraining stops being a liability and becomes a clean low-label lever: the full framework EquiSSL (GE-RPE $+$ $\pi_R$ $+$ iBOT$+$MAE) tightens the drop to $0.8\%$ at $68.30\%$ val mIoU and lifts $1\%$-label fine-tuning by $+2.39$ mIoU on the $N{=}373$ test split (and by $+4.10$ on the smaller $N{=}40$ val split); the same fix carries over to monocular depth and to zero-shot Structured3D segmentation. The construction is provably $C_n$-invariant and $\mathcal{O}(n^{-2})$-close to the continuous $\mathrm{SO}(2)$ average, making the resulting model a usable $360^\circ$ visual-computing primitive across panoramic relighting, immersive video, and cross-dataset transfer. Code is available at https://github.com/Jaywalk18/equissl-release.

Original source

This story was published by arXiv cs.CV and written by Tianjian Zhou, Yishan Li, Jie Jiang, Yifei Zhang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News