
QY
Qilang Ye, Meng Liu, Yu Zhou
· 1 min read
ResearcharXiv cs.AI
RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation
arXiv:2609.32224v1 Announce Type: new
Abstract: We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.
Original source
This story was published by arXiv cs.AI and written by Qilang Ye, Meng Liu, Yu Zhou. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


