Zero-Shot Affordance Exposure for Robotic Manipulation Via Object Repositioning
Abstract
Robotic manipulation relies on perceiving the object region that affords an instructed interaction, yet often assumes that this region is visible and accessible. In cluttered scenes, however, a handle, blade, tip, or other task-relevant region may be occluded. This paper introduces a zero-shot affordance exposure framework that couples language-conditioned visual perception with geometrically grounded pick-and-place actions. A visionlanguage model (VLM) uses a task instruction and an RGB observation to identify the required interaction region, assess its exposure, and produce structured scene-revision parameters. Segment Anything supplies a manipulation mask, while a zeroshot geometric grasp generator produces an executable top-down grasp within the masked safe region for a Franka Research 3 arm. The robot repositions an occluding object to reveal the required region; if it is already exposed, the target is grasped directly at a safe region. Real-robot experiments with knife, scissors, and screwdriver scenes show that the framework exposes taskrelevant regions under different occlusion conditions and directly grasps targets when scene repair is unnecessary, without taskspecific training. This perception-action loop provides a practical pre-manipulation mechanism for cluttered tabletop tasks while grounding robot execution in geometric visual evidence.