
JK
Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang
· 1 min read
ResearcharXiv cs.CV
Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning
arXiv:2609.20235v1 Announce Type: new
Abstract: Indoor mobile robots require open-vocabulary scene understanding that grounds natural-language queries in a consistent 3D map. Many existing systems ultimately rely on cosine-similarity retrieval with contrastive image--text encoders, which is efficient but brittle when labels are near-synonymous or multiple similar instances appear. We present Scene-Q, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases. High-confidence queries are answered by fast retrieval, while ambiguous ones are reranked over a small top-K candidate set using the original multi-view images and instance bounding boxes, enabling context-aware disambiguation at low cost. Scene-Q improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real-world reconstructions, with the largest gains on spatial and relational queries while keeping a substantial fraction of queries on the fast path.
Original source
This story was published by arXiv cs.CV and written by Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


