
AR
Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard
· 1 min read
ResearcharXiv cs.CV
One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation
arXiv:2603.23488v3 Announce Type: replace
Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training to a narrow set of purpose-built datasets. We propose in-the-wild monocular pretraining: a frozen depth estimator lifts each source image into 3D and reprojects under sampled poses to yield pseudo-target views; masked losses restrict supervision to valid regions and an adversarial objective covers disoccluded areas. Scaled to 30 million uncurated images, this produces OVIE, requiring only a source image and target pose at inference. Prior work trains without multi-view data but needs a depth estimator at inference, or drops this dependency but requires multi-view training pairs; OVIE is the first to require neither. Without multi-view supervision, OVIE rivals in-domain baselines on RealEstate10K and surpasses all on DL3DV, producing the most multi-view-consistent trajectories of any geometry-free method; brief multi-view fine-tuning outperforms all geometry-free methods on their training domain. At 116 FPS, it is over 600x faster than the fastest baseline. Code and pretrained models are at https://github.com/kyutai-labs/ovie; video results are on the project page, https://kyutai.org/blog/2026-04-14-ovie/.
Original source
This story was published by arXiv cs.CV and written by Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


