Skip to content

Author

Haibin Huang

We have 8 of 33 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world models are largely driven by appearance priors and often lose physical or spatial consistency over long h...

Xin Zhang, Yabo Chen, Zi-Xuan Duan et al. · 2 citations
Preprint Jul 2026

Generative Transmission: Rethinking Computation, Bandwidth, and Memory in Communication

Under the AI Flow framework, communication is shifting from transmitting fidelity-oriented information flows toward delivering task-oriented and perception-oriented token flows across heterogeneous network resources. Video communication is a fundamental component of modern information networks. However, under ultra-low...

Xiangyu Chen, Jixiang Luo, Yuan-Kai Fan et al. · 2 citations
Jul 2026

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately a...

Yuyang Huang, Yabo Chen, Wenrui Dai et al. · 4 citations
Preprint Aug 2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples, is introduced and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency.

Bojia Zi, Xiaoyan Yang, Yu Zhou et al. · 0 citations
Preprint Aug 2026

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

A unified latent-space framework for image and video diffusion models that achieves the sota performance among various metrics and further improves optimization stability and achieves the highest VBench quality, semantic, and total scores among the evaluated methods.

Rui Li, Yuan-Zhi Liang, Ke-Chun Hao et al. · 0 citations
Review Jul 2026

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

This work proposes a co-evolution roadmap for physical intelligence centered on theembodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.

Yuanzhi Liang, Xufeng Zhan, Haibin Huang et al. · 1 citation
Jun 2026

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.

Haoyuan Wang, Yabo Chen, Haibin Huang et al. · 1 citation
Jun 2026

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure,"motion entanglement", the unintended coupling of independent motion sources, such as camera movement and object motion. We identify tha...

Ruoyu Wang, Jialun Liu, Huayang Huang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.