
JL
Junming Liu, Yuqi Li, Yifei Sun, Maonan Wang, Yiming Cheng, Rui Qian, Piotr Koniusz, Yirong Chen, Ding Wang
· 1 min read
ResearcharXiv cs.AI
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
arXiv:2605.18162v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) have made striking progress, yet their spatial reasoning remains fragile. Models that answer an original input correctly can still fail under valid transformations with predictable answer mappings, revealing a gap between instance-level correctness and robust spatial reasoning. To address this, we propose Spatial Alignment via Geometric Evolution (SAGE), a self-evolving framework that improves robust spatial reasoning through geometric and linguistic duality operations. SAGE incorporates duality consistency into GRPO training, encouraging models to produce coherent answers across original and transformed inputs. SAGE co-evolves duality generation and solution, allowing the model to continually expose and address its own reasoning weaknesses. A dynamic operation pool identifies challenging operations and retires mastered ones, keeping training focused on informative duality signals. SAGE is model-agnostic, data-efficient compared to prior post-training methods, and can be applied as a lightweight adaptation stage to any existing VLM. Experiments on video and spatial reasoning benchmarks demonstrate consistent improvements over strong baselines and enhanced generalization to unseen data.
Original source
This story was published by arXiv cs.AI and written by Junming Liu, Yuqi Li, Yifei Sun, Maonan Wang, Yiming Cheng, Rui Qian, Piotr Koniusz, Yirong Chen, Ding Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


