Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world models are largely driven by appearance priors and often lose physical or spatial consistency over long h...
Xin Zhang, Yabo Chen, Zi-Xuan Duan et al.· 2 citations
Robust navigation in cluttered environments remains a fundamental challenge for quadrotors, particularly when strong wind disturbances arise, which perturb vehicle dynamics, limit control authority, and substantially increase collision risk. Existing learning-based navigation policies typically rely on obstacle percept...
Zhonghan Tang, Chenhui Li, Shuai Liang et al.· IEEE Robotics and Automation...· 0 citations
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately a...
Yuyang Huang, Yabo Chen, Wenrui Dai et al.· arXiv.org· 4 citations
WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.
Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al.· 1 citation
RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples, is introduced and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency.
Bojia Zi, Xiaoyan Yang, Yu Zhou et al.· 0 citations
An A2I model, AudioCanvas, fine-tuned on the A2I-Set is proposed, a unified, high-quality tri-modal dataset specifically designed for audio-visual research, including audio-conditioned image generation.
Dongxu Ge, Shansong Liu, Cheng Gong et al.· 0 citations
This work proposes a co-evolution roadmap for physical intelligence centered on theembodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.
Yuanzhi Liang, Xufeng Zhan, Haibin Huang et al.· 1 citation
This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.
Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure,"motion entanglement", the unintended coupling of independent motion sources, such as camera movement and object motion. We identify tha...