Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world models are largely driven by appearance priors and often lose physical or spatial consistency over long h...
Xin Zhang, Yabo Chen, Zi-Xuan Duan et al.· 2 citations
Under the AI Flow framework, communication is shifting from transmitting fidelity-oriented information flows toward delivering task-oriented and perception-oriented token flows across heterogeneous network resources. Video communication is a fundamental component of modern information networks. However, under ultra-low...
Xiangyu Chen, Jixiang Luo, Yuan-Kai Fan et al.· 2 citations
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately a...
Yuyang Huang, Yabo Chen, Wenrui Dai et al.· arXiv.org· 4 citations
RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples, is introduced and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency.
Bojia Zi, Xiaoyan Yang, Yu Zhou et al.· 0 citations
A unified latent-space framework for image and video diffusion models that achieves the sota performance among various metrics and further improves optimization stability and achieves the highest VBench quality, semantic, and total scores among the evaluated methods.
Rui Li, Yuan-Zhi Liang, Ke-Chun Hao et al.· 0 citations
This work proposes a co-evolution roadmap for physical intelligence centered on theembodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.
Yuanzhi Liang, Xufeng Zhan, Haibin Huang et al.· 1 citation
This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.
Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure,"motion entanglement", the unintended coupling of independent motion sources, such as camera movement and object motion. We identify tha...