AVLF-Nav: Embodied Object Navigation via Active Visual Exploration and Lightweight Vision–Language Fusion
Abstract
Embodied object navigation must couple partial visual perception with efficient exploration in unseen indoor environments. Existing agents often align language and vision without explicitly deciding which viewpoint is most informative, causing repeated visits and inefficient trajectories. This paper proposes AVLF-Nav, a lightweight framework integrating target-guided cross-modal attention, an online exploration graph, graph-attention reasoning, and active candidate-viewpoint scoring. The utility function balances semantic relevance, unknown-area gain, reachability, travel cost, and revisit frequency before a gated recurrent policy predicts discrete actions. Experiments in an indoor-scene-based simulated environment show a 72.3% success rate and 0.58 SPL, exceeding the strongest graph baseline by 5.9 percentage points and 0.06. Ablation, noise, occlusion, low-data, and case-level analyses verify improved efficiency, robustness, and interpretability.