Vision-language models (VLMs) show strong visual understanding for aerial navigation, but their action generation remains unreliable. We study this gap in the context of orbit-search-based navigation, where an aerial agent circles a reference landmark while a VLM detects a language-described goal. We find that the VLM...
Hao-Tian Xu, Chen-Xu Wang, De-Jun Chen et al.· 2026 12th International Conf...· 0 citations
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth...
SBFNav is introduced, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF) that preserves multiple spatial hypotheses under par- tial evidence and further confirms the advantages of spatial-belief modeling over single-point prediction.
Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.