MMIRL: a multimodal framework for learning robotic manipulation from RGB-D demonstrations
Abstract
Extracting robot-reproducible manipulation skills from human RGB-D demonstrations requires stable object identities under multi-object occlusion and a conversion from continuous observations to structured demonstrations. MMIRL addresses this setting with a modular two-stage framework for tabletop manipulation. Offline, the method generates paired synthetic data from a Unified Robot Description Format (URDF) object library and combines it with mixed synthetic data rendered on real backgrounds to calibrate segmentation prompts, candidate selection, and cross-frame association. Online, a detector provides box prompts, a promptable segmentation model predicts instance masks, and identity maintenance performs one-to-one association under joint appearance and geometric constraints to recover identity-aware instance trajectories and 2.5D states. Hand keypoint trajectories then localize interaction intervals and manipulated objects, enabling object-relative normalized snippets for conditional skill and trajectory learning. Evaluations on paired synthetic data and real demonstrations show stable cross-frame association under varying object counts and occlusion, and yield reliable structured demonstrations for replay and execution validation.