This work shows that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass, enabling interactive 4D-controllable streaming generation for the first time.
Shiqian Li, Chenguo Lin, Zhi-Guang Liu et al.· 0 citations
A Aura, a unified framework for high-fidelity and identity-consistent video generation, and introduces AI director-level captions that provide dense and structured descriptions of video content to better capture scene dynamics and subject interactions.
Zixiang Zhou, Zhentao Yu, Yifeng Ma et al.· 0 citations