Skip to content
Conference

Semantic relevance guided grounding for MLLM-based embodied navigation

Aug 2026 · International Conference on Advanced Sensing and Intelligent Systems · Vol 14309, pp. 1430916 - 1430916-6 · 0 citations · 14 references
Engineering

Abstract

Multimodal Large Language Models (MLLMs) based Embodied navigation faces a severe challenge where key cues are easily overwhelmed by complex environmental noise, leading to inefficient decision-making. To address this, we propose a Semantic Relevance Guided grounding enhanced navigation framework(SRG-Nav). The core idea of our approach lies in utilizing semantic relevance to guide visual and language attention. By evaluating the correlation between scene entities and the navigation goal, SRG-Nav adds ranked high-value cues to system prompts and maps them back into the visual space to generate explicit bounding boxes. This mechanism explicitly directs the MLLM to focus on task-relevant entities and regions while effectively suppressing environmental noise. Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.

View source