SyncAI.news, a Varaisys broadcasting
Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding
BP

Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang

· 1 min read

ResearcharXiv cs.CV

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

arXiv:2609.29835v1 Announce Type: new Abstract: LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.

Original source

This story was published by arXiv cs.CV and written by Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News