SyncAI.news, a Varaisys broadcasting
Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation
AF

Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua

· 1 min read

ResearcharXiv cs.CV

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

arXiv:2609.24850v1 Announce Type: new Abstract: Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.

Original source

This story was published by arXiv cs.CV and written by Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News