Skip to content
Preprint

Map2Route: Benchmarking Compositional Language-Grounded Route Planning over Semantic Maps

Sep 2026 · 0 citations · 22 references
Computer Science

TL;DR

Across seven representative adapted baselines, Grounding2Route substantially outperforms existing methods in all metrics, and a substantial gap to human demonstrations remains, highlighting the difficulty of Map2Route and the considerable headroom for future progress.

Abstract

We introduce Map2Route, a human-curated benchmark for compositional language-grounded route planning over pre-built semantic maps. Map2Route contains 1,000 episodes across 40 scenes, where instructions use relational, comparative, and nested descriptions to identify route-relevant objects and regions, while specifying ordered must-pass regions, must-avoid requirements, five categories of soft preferences, and spatial and route-stage scopes, which is partially tested by existing works. Alongside Map2Route, we propose Grounding2Route, which combines executable code-as-grounding with verification-guided repair and scope-aware planning.Across seven representative adapted baselines, Grounding2Route substantially outperforms existing methods in all metrics. Despite these gains, a substantial gap to human demonstrations remains, highlighting the difficulty of Map2Route and the considerable headroom for future progress. Additional qualitative results and resources are available on https://anonymous.4open.science/w/Map2Route-F05F/.

View source

Similar papers

ZIVIL: Zero-Shot Incremental Vision-Language Maps and Spatial Graph Representation of Construction Sites

This framework combines simultaneous localization and mapping (SLAM), visual-language feature extraction, incremental semantic and instance label fusion, and spatial graph construction to enable a construction robot navigation framework that supports open-vocabulary language queries.

Charles M. Raines, I. Fernandez, Mandy Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience

Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instructi...

Si-Qi Zhang, Meng Wei, Chen-Yang Wan et al. · 0 citations
Preprint Aug 2026

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

This work introduces SPATIALQUERY, a training- free framework for CIDQ reasoning from a single RGB image, together with SPATIALQUERY-1M, a benchmark containing over one million RGB-only question-answer pairs from 200 indoor scenes, and proposes Uncertainty-Aware Chain-of-Thought (UA-CoT) prompting, which incorporates g...

Hai-Tra Nguyen, Tung Vu, Cong Tran · 0 citations
Preprint Sep 2026

Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how alig...

Xun Huang, Shi-Jia Zhao, Rong Qu et al. · 0 citations
Preprint Sep 2026

GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation

Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We in...

Kang-Ping Luo, He-Sheng Wang · 0 citations
Preprint Sep 2026

NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route

Human-readable maps provide an intuitive interface for specifying robot destinations, but connecting their schematic geometry to egocentric observations remains challenging. We present NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrai...

Ayun Lee, Jiseon Kim, Giseop Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.