Skip to content
Conference

MMIRL: a multimodal framework for learning robotic manipulation from RGB-D demonstrations

Jul 2026 · International Conference on Hydromechatronics and Advanced Robot Control Technology · Vol 14253, pp. 142530R - 142530R-8 · 0 citations · 6 references
Engineering

Abstract

Extracting robot-reproducible manipulation skills from human RGB-D demonstrations requires stable object identities under multi-object occlusion and a conversion from continuous observations to structured demonstrations. MMIRL addresses this setting with a modular two-stage framework for tabletop manipulation. Offline, the method generates paired synthetic data from a Unified Robot Description Format (URDF) object library and combines it with mixed synthetic data rendered on real backgrounds to calibrate segmentation prompts, candidate selection, and cross-frame association. Online, a detector provides box prompts, a promptable segmentation model predicts instance masks, and identity maintenance performs one-to-one association under joint appearance and geometric constraints to recover identity-aware instance trajectories and 2.5D states. Hand keypoint trajectories then localize interaction intervals and manipulated objects, enabling object-relative normalized snippets for conditional skill and trajectory learning. Evaluations on paired synthetic data and real demonstrations show stable cross-frame association under varying object counts and occlusion, and yield reliable structured demonstrations for replay and execution validation.

View source