Skip to content
Open access

Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes

Sep 2026 · Robotics · 0 citations · 6 references

Abstract

Retrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide model fine-tuning and evaluation. In addition, comprehensive evaluation requires accounting for the full space of topologically valid clearing sequences while systematically disentangling high-level topological planning errors from low-level physical execution failures. To fulfill these requirements, we introduce a novel DAG-based obstruction reasoning framework coupled with an integrated diagnostic evaluation protocol. Specifically, we model scene-level physical dependencies as Directed Acyclic Graphs (DAGs), explicitly converting graph constraints into topologically feasible sequence permutations to drive VLM fine-tuning. For diagnostic evaluation, we establish a three-part offline protocol comprising Strict Exact Match (EM), Graph-Feasible Accuracy (GFA), and Multi-Reference Normalized Sequence Edit Distance (MR-NSED), paired with online simulation testing. Extensive experiments demonstrate that our topological fine-tuning significantly improves multi-path reasoning performance, outperforming strong foundation model baselines including GPT-4o, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct, while our diagnostic protocol provides a faithful mechanism to systematically isolate reasoning logic from manipulation mechanics in complex physical clutter.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.