SyncAI.news, a Varaisys broadcasting
REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
SA

Sheir A. Zaheer, Jihwan Moon, Chan Y. Park

· 1 min read

ResearcharXiv cs.CV

REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

arXiv:2610.07585v1 Announce Type: new Abstract: We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.

Original source

This story was published by arXiv cs.CV and written by Sheir A. Zaheer, Jihwan Moon, Chan Y. Park. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News