SyncAI.news, a Varaisys broadcasting
FAST: Flow Any Scene Transformer
YZ

Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo

· 1 min read

ResearcharXiv cs.CV

FAST: Flow Any Scene Transformer

arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.

Original source

This story was published by arXiv cs.CV and written by Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News