Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragi...
Wei-Lun Feng, Chuan-Guang Yang, Haotong Qin et al.· IEEE Transactions on Pattern...· 3 citations
Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorl...
Yuqi Li, Xi Xiao, Yunbei Zhang et al.· arXiv.org· 4 citations
3D scene generation has rapidly evolved, significantly promoting the innovation of content creation. In this context, interaction techniques serve as a pivotal bridge connecting user intent with the generative models, thereby enabling precise control, real-time feedback and personalized customization of complex 3D scen...
Yuqi Li, Si-Wei Meng, Chuan-Guang Yang et al.· Proceedings of the Thirty-Fi...· 30 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.