Skip to content
Open access

Hardware-Aware Acceleration of Open-Vocabulary Multi-Object Navigation on Edge GPUs

Sep 2026 · Electronics · 0 citations

Abstract

Persistent open-vocabulary navigation enables robots to search for sequential language-specified objects using reusable visual semantic evidence. On edge GPUs, the pipeline is constrained by foundation-model inference, semantic projection, persistent map movement, frontier processing, and repeated target detection. Using OneMap as a controlled reference, we reorganize this workload through compiled mixed-precision perception, direct-to-map confidence-weighted feature aggregation, accelerator-resident semantic maps, compiled navigation operations, and language-conditioned detector scheduling. The resulting dataflow removes the dense image resolution feature intermediate and schedules target confirmation detection from current-view target similarity. On the complete Habitat-Matterport 3D benchmarks, our implementation retains 92.6% and 96.1% of the reproduced OneMap single- and multi-object success rates. Across 36 paired hardware-in-the-loop comparisons on a Jetson AGX Orin at four power modes, it achieves geometric mean speedups of 4.64× for on-device computation and 3.55× for the complete episode duration. It also reduces peak system random access memory (RAM) by 51.21% and a summed board-rail energy proxy by 75.32% on average while preserving every paired outcome. These results establish the coordinated dataflow and current-view target-conditioned execution as effective mechanisms for efficient persistent foundation-model navigation on power-constrained edge robots.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.