Skip to content
Conference Open access

VLEG: Embodied Vision-Language Grasping for a Quadruped Manipulator

Aug 2026 · Journal of Physics, Conference Series · Vol 3291 · 0 citations · 37 references
Physics

Abstract

Open-vocabulary grasping on a quadruped manipulator requires more than recognizing the target object. The robot must also select a grasp pose that is both consistent with the task semantics and reliable to execute under body motion and viewpoint changes. In this paper, we present VLEG, an embodied vision-language grasping framework for quadruped manipulators that explicitly incorporates body motion into grasp decision making. Our method guides the robot to continuously adjust its body pose during approach and optimize local observations before grasping, thereby improving perception quality. For grasp decision making, instead of using a coarse single-stage filtering strategy, we design a multi-stage and multi-criteria grasp selection mechanism based on geometric grasp candidates. This mechanism jointly considers physical feasibility and task consistency. We implement the complete system on an onboard Jetson platform and conduct extensive real-world experiments on a quadruped robot equipped with a manipulator, covering tabletop, low-platform, ground-level, and outdoor raised-platform scenes. The results validate the deployability of VLEG in real-world quadruped manipulation scenarios, as well as its robust grasping ability and task-aware decision-making capability across the tested object categories.

Read PDF