Jul 2026
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
DeVA is introduced, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance, enabling rich information exchange while making policy learning more tractable.
Mengqi Zhang, Sahil Khose, Simar Kareer et al.
· arXiv.org · 2 citations