
HH
Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li, Maosong Sun, Ruixuan Li
· 1 min read
ResearcharXiv cs.CV
MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding
arXiv:2609.37374v1 Announce Type: new
Abstract: Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.
Original source
This story was published by arXiv cs.CV and written by Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li, Maosong Sun, Ruixuan Li. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


