Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, t...
Xin-Yi Chen, Han-Xin Zhu, Xi-Jun Wang et al.· 0 citations
Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explic...
Sen Liang, Fengbin Guan, Youliang Zhang et al.· 5 citations
This work presents Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis, and proposes OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding.
Liang Xu, Cheng-Qun Yang, Zili Lin et al.· 0 citations
IMPACT is introduced, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting, which consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
Rong-Ze Tang, Jianjie Fang, Zhao-Lu Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.