Preprint
Jul 2026
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
A Aura, a unified framework for high-fidelity and identity-consistent video generation, and introduces AI director-level captions that provide dense and structured descriptions of video content to better capture scene dynamics and subject interactions.
Zixiang Zhou, Zhentao Yu, Yifeng Ma et al.
· 0 citations