This paper provides a formalization of what constitutes a sufficient scene graph for planning by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation.
Abstract
Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general, task-centric formalization of what constitutes a sufficient scene graph for planning remains open. This paper provides such a formalization by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation. We then introduce derived scene graphs via information mappings that merge and prune nodes and induce quotient transition systems augmented with motion primitives to capture higher-level actions over merged graph nodes. Sufficiency is characterized by two conditions: (i) the information mapping yields a deterministic quotient, and (ii) the task is well-posed over derived traces, ensuring plans found on the derived model are feasible on the maximal system. We illustrate the framework using a task over an example environment, showing both sufficient and insufficient reduced scene graphs.
GRAB-TAMP is introduced, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency.
N. Vijayakumar, Nav Singhal, G. Varma et al.· 0 citations
3D scene graphs provide semantically rich and hierarchical representations for robot perception. However, existing systems do not maintain uncertainty as an explicit belief or propagate it through the operations that construct and refine the graph. We introduce Probabilistic Scene Graph (PSG), a generalization of the c...
Waqas Ali, Michele Antonazzi, Timon Homberger et al.· 0 citations
Experimental results demonstrate that current methods struggle to complete the ESRP task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning.
Can-Zhi Chen, Zan Wang, Si-Qi Zhu et al.· IEEE Robotics and Automation...· 0 citations
An adaptive task planning method based on video priors and dynamic scene graphs (ATP-VPDSG) that leverages the VLM to extract manipulation logic from video demonstrations, thus supplementing manipulation priors and outperforming the selected task planning baselines.
Guang-Hui Ma, Jia-Hui Guo, Xin-Hua Tang et al.· Italian National Conference...· 0 citations
Autonomous navigation across heterogeneous environments remains challenging because different environments may require fundamentally different planning representations and action models. Structured roads are naturally represented as graphs with constrained connectivity, unstructured terrain requires continuous geometri...