Preprint
Aug 2026
SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control
SelfWAM is introduced, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion.
Bikang Pan, Fan Liu, Haotao Lu et al.
· 1 citation