Skip to content

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Jul 2026 · arXiv.org · Vol abs/2607.24159 · 2 citations · 67 references
Computer Science

TL;DR

DeVA is introduced, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance, enabling rich information exchange while making policy learning more tractable.

Abstract

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.

View source

Similar papers

Preprint Sep 2026

XPACE: Joint World and Action Modeling from Heterogeneous Experience

A general-purpose robot needs to draw on diverse experience, choose actions, and anticipate how those actions will change the world. We introduce XPACE, a unified embodied world model that serves as both a world action model, jointly predicting executable robot actions and future video, and a world simulator, predictin...

Jiacheng Wei, J. Bai, X. Yue et al. · 1 citation
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 4 citations
Preprint Aug 2026

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Vid2WAM is proposed, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student and introduces source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from nois...

Chen-Hao Qiu, Ruixiang Wang, Runyi Zhao et al. · 1 citation
Preprint Sep 2026

SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models

Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive i...

Tian-Fu Li, Hao-Xuan Xu, Wen-Bo Chen et al. · 0 citations
Preprint Aug 2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

World Tokens is an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation and is highly competitive on LIBERO, attains the best reported averages on SIMPLER, and substantially improves real-world R1 Pro success over a matched...

Qu Tang, Benhui Zhuang, Bo Yuan et al. · 1 citation · ⚡1
Preprint Sep 2026

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...

Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.