Sep 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 39 references
Medicine
TL;DR
An adaptive task planning method based on video priors and dynamic scene graphs (ATP-VPDSG) that leverages the VLM to extract manipulation logic from video demonstrations, thus supplementing manipulation priors and outperforming the selected task planning baselines.
Abstract
Robots are now expected to execute increasingly complex long-horizon tasks in unstructured environments. Despite the strong potential of pretrained Vision-Language Models (VLMs) in task planning, their direct application to robotic manipulation is hindered by logical reasoning deviations and inadequate geometric scene perception. This work proposes an adaptive task planning method based on video priors and dynamic scene graphs (ATP-VPDSG). It leverages the VLM to extract manipulation logic from video demonstrations, thus supplementing manipulation priors. Meanwhile, scene graphs were integrated to convert unstructured environments into structured representations with spatial topological relations, compensating for perceptual deficiencies. A dual-track feedback mechanism based on visual expectations was further incorporated to enable failure diagnosis and adaptive replanning in complex environments. Extensive long-horizon robotic manipulation experiments were conducted on the LIBERO-10 benchmark with Qwen3-VL as the core VLM. Results showed that ATP-VPDSG achieved an average task planning accuracy of 91.2% and a task execution success rate of 74.67%, outperforming the selected task planning baselines. Ablation studies verified that video priors and dynamic scene graphs exerted complementary effects on logical constraints and physical feasibility. Furthermore, a real-robot experiment on an industrial slider–rail assembly task demonstrated successful sim-to-real transfer, achieving an 82.0% success rate without task-specific fine-tuning.
Robotic-GST is presented, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment.
Si-Chao Liu, Ze-Kun Wang, Li-Xuan Tang et al.· 0 citations
This paper provides a formalization of what constitutes a sufficient scene graph for planning by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation.
TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
Evidence Acquisition and Feasibility Gating (EAFG) is proposed, a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution and applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt.
Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al.· 0 citations
Results show that the Simify framework outperforms prior work on foundation models for spatial reasoning by effectively exploiting large-scale parallel simulation during inference, and also highlight the importance of complete and accurate geometry for successful sim-to-real transfer.
Ivan Kapelyukh, Ya-Fei Hu, Ran Gong et al.· 0 citations
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-a...
Ho-Jin Lee, S. Li, Maximilian Hilger et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.