Skip to content

Author

Xuelong Li

We have 9 of 53 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world models are largely driven by appearance priors and often lose physical or spatial consistency over long h...

Xin Zhang, Yabo Chen, Zi-Xuan Duan et al. · 2 citations
Nov 2026

WAND: Learning Robust Navigation Under Complex Wind Disturbances and Dense Obstacles for Quadrotors

Robust navigation in cluttered environments remains a fundamental challenge for quadrotors, particularly when strong wind disturbances arise, which perturb vehicle dynamics, limit control authority, and substantially increase collision risk. Existing learning-based navigation policies typically rely on obstacle percept...

Zhonghan Tang, Chenhui Li, Shuai Liang et al. · 0 citations
Jul 2026

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately a...

Yuyang Huang, Yabo Chen, Wenrui Dai et al. · 4 citations
Preprint Aug 2026

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.

Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al. · 1 citation
Preprint Aug 2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples, is introduced and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency.

Bojia Zi, Xiaoyan Yang, Yu Zhou et al. · 0 citations
Review Jul 2026

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

This work proposes a co-evolution roadmap for physical intelligence centered on theembodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.

Yuanzhi Liang, Xufeng Zhan, Haibin Huang et al. · 1 citation
Jun 2026

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.

Haoyuan Wang, Yabo Chen, Haibin Huang et al. · 1 citation
Jun 2026

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure,"motion entanglement", the unintended coupling of independent motion sources, such as camera movement and object motion. We identify tha...

Ruoyu Wang, Jialun Liu, Huayang Huang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.